Conversation
|
This duplicates a bunch of the implementation over in #3405, so will make some suggestions/comments on that one and then rebase this once that PR lands |
|
I guess we should coordinate more as I also was preparing PR for std_normal and normal |
Jenkins Console Log Machine informationDistributor ID: Ubuntu Description: Ubuntu 20.04.3 LTS Release: 20.04 Codename: focal CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Byte Order: Little Endian Address sizes: 52 bits physical, 57 bits virtual CPU(s): 192 On-line CPU(s) list: 0-191 Thread(s) per core: 2 Core(s) per socket: 48 Socket(s): 2 NUMA node(s): 2 Vendor ID: AuthenticAMD CPU family: 25 Model: 17 Model name: AMD EPYC 9474F 48-Core Processor Stepping: 1 Frequency boost: enabled CPU MHz: 1497.515 CPU max MHz: 4114.4229 CPU min MHz: 1500.0000 BogoMIPS: 7189.39 Virtualization: AMD-V L1d cache: 3 MiB L1i cache: 3 MiB L2 cache: 96 MiB L3 cache: 512 MiB NUMA node0 CPU(s): 0-47,96-143 NUMA node1 CPU(s): 48-95,144-191 Vulnerability Gather data sampling: Not affected Vulnerability Indirect target selection: Not affected Vulnerability Itlb multihit: Not affected Vulnerability L1tf: Not affected Vulnerability Mds: Not affected Vulnerability Meltdown: Not affected Vulnerability Mmio stale data: Not affected Vulnerability Old microcode: Not affected Vulnerability Reg file data sampling: Not affected Vulnerability Retbleed: Not affected Vulnerability Spec rstack overflow: Mitigation; Safe RET Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; STIBP always-on; PBRSB-eIBRS Not affected; BHI Not affected Vulnerability Srbds: Not affected Vulnerability Tsa: Mitigation; Clear CPU buffers Vulnerability Tsx async abort: Not affected Vulnerability Vmscape: Mitigation; IBPB before exit to userspace Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good amd_lbr_v2 nopl xtopology nonstop_tsc cpuid extd_apicid aperfmperf rapl pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt tce topoext perfctr_core perfctr_nb bpext perfctr_llc mwaitx cpb cat_l3 cdp_l3 hw_pstate ssbd mba perfmon_v2 ibrs ibpb stibp ibrs_enhanced vmmcall fsgsbase bmi1 avx2 smep bmi2 erms invpcid cqm rdt_a avx512f avx512dq rdseed adx smap avx512ifma clflushopt clwb avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local user_shstk avx512_bf16 clzero irperf xsaveerptr rdpru wbnoinvd amd_ppin cppc arat npt lbrv svm_lock nrip_save tsc_scale vmcb_clean flushbyasid decodeassists pausefilter pfthreshold avic v_vmsave_vmload vgif x2avic v_spec_ctrl vnmi avx512vbmi umip pku ospke avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg avx512_vpopcntdq la57 rdpid overflow_recov succor smca fsrm flush_l1d debug_swap G++: g++ (Ubuntu 9.4.0-1ubuntu1~20.04) 9.4.0 Copyright (C) 2019 Free Software Foundation, Inc. This is free software; see the source for copying conditions. There is NO warranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. Clang: clang version 10.0.0-4ubuntu1 Target: x86_64-pc-linux-gnu Thread model: posix InstalledDir: /usr/bin |
|
I reviewed this PR (with help from Claude). Looks great Accuracy against mpmathWorst ulp over 601 points per range:
For comparison, develop's worst in-range relative error is 7.6e-06 to 5.6e-05, which is 3.4e+10 to 2.5e+11 ulp, across every branch except the far Mills tail. So this PR improves the gradient by about ten orders of magnitude. One suggestion: recover the rounding error of
|
| range | value now | value with fma | gradient now | gradient with fma |
|---|---|---|---|---|
| [0.6, 4] | 6.7 | 4.8 | 4.8 | 2.2 |
| [4, 10] | 30.1 | 3.5 | 30.3 | 2.6 |
| [10, 20] | 115.2 | 2.8 | 120.9 | 2.7 |
| [20, 37] | 494.4 | 3.1 | 503.2 | 2.2 |
The cost is 3 to 4 percent on a Xeon E5-2680 v3 and 2 to 5 percent on a Tesla V100, measured with the value and gradient computed together.
Two notes on this. It stays applicable after the rebase onto #3405, because it concerns the exp(-z^2/2) factor and not erfcx.
The same expression appears in std_normal_lcdf_impl in the OpenCL device function, so the change belongs in both places.
erfcx looks slightly faster on the lower tail
Beyond removing the duplicated Cody coefficients, #3405 may improve speed. Comparing your internal std_normal_erfcx path against calling erfcx directly, value and gradient computed together:
| range (scaled) | CPU | GPU |
|---|---|---|
| far tail [-40, -10] | 1.16 | 1.35 |
| Cody [-10, -4] | 1.10 | 2.21 |
| Cody edge [-4, -2.5] | 1.06 | 1.03 |
| interior [-2.5, -0.1] | 1.11 | 1.93 |
Ratios above 1 mean the direct erfcx call is faster.
|
awesome! thanks @andrjohns for putting this together, it's much needed |
Two more functions get the benefit with no extra workTwo functions that benefit from this
One function benefits from erfcx directly and is not covered here
return std_normal_lcdf(-x) - std_normal_lpdf(x);It calls
The loss grows without bound. |
|
@andrjohns this is very cool! Because of how tied in this and #3405 are do you want to handle the #3405 review and then I'll review this once that is in? |
Yep sounds like a plan to me! |
|
@andrjohns ping me when this is ready for review |
|
@andrjohns is this ready for me to look at? |
Jenkins Console Log Machine informationDistributor ID: Ubuntu Description: Ubuntu 20.04.3 LTS Release: 20.04 Codename: focal CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Byte Order: Little Endian Address sizes: 43 bits physical, 48 bits virtual CPU(s): 256 On-line CPU(s) list: 0-255 Thread(s) per core: 2 Core(s) per socket: 64 Socket(s): 2 NUMA node(s): 2 Vendor ID: AuthenticAMD CPU family: 23 Model: 49 Model name: AMD EPYC 7742 64-Core Processor Stepping: 0 Frequency boost: enabled CPU MHz: 1496.992 CPU max MHz: 3416.0681 CPU min MHz: 1500.0000 BogoMIPS: 4491.85 Virtualization: AMD-V L1d cache: 4 MiB L1i cache: 4 MiB L2 cache: 64 MiB L3 cache: 512 MiB NUMA node0 CPU(s): 0-63,128-191 NUMA node1 CPU(s): 64-127,192-255 Vulnerability Gather data sampling: Not affected Vulnerability Indirect target selection: Not affected Vulnerability Itlb multihit: Not affected Vulnerability L1tf: Not affected Vulnerability Mds: Not affected Vulnerability Meltdown: Not affected Vulnerability Mmio stale data: Not affected Vulnerability Old microcode: Not affected Vulnerability Reg file data sampling: Not affected Vulnerability Retbleed: Mitigation; untrained return thunk; SMT enabled with STIBP protection Vulnerability Spec rstack overflow: Mitigation; Safe RET Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Vulnerability Spectre v2: Mitigation; Retpolines; IBPB conditional; STIBP always-on; RSB filling; PBRSB-eIBRS Not affected; BHI Not affected Vulnerability Srbds: Not affected Vulnerability Tsa: Not affected Vulnerability Tsx async abort: Not affected Vulnerability Vmscape: Mitigation; IBPB before exit to userspace Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good nopl xtopology nonstop_tsc cpuid extd_apicid aperfmperf rapl pni pclmulqdq monitor ssse3 fma cx16 sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt tce topoext perfctr_core perfctr_nb bpext perfctr_llc mwaitx cpb cat_l3 cdp_l3 hw_pstate ssbd mba ibrs ibpb stibp vmmcall fsgsbase bmi1 avx2 smep bmi2 cqm rdt_a rdseed adx smap clflushopt clwb sha_ni xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local clzero irperf xsaveerptr rdpru wbnoinvd amd_ppin arat npt lbrv svm_lock nrip_save tsc_scale vmcb_clean flushbyasid decodeassists pausefilter pfthreshold avic v_vmsave_vmload vgif v_spec_ctrl umip rdpid overflow_recov succor smca sev sev_es G++: g++ (Ubuntu 9.4.0-1ubuntu1~20.04) 9.4.0 Copyright (C) 2019 Free Software Foundation, Inc. This is free software; see the source for copying conditions. There is NO warranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. Clang: clang version 10.0.0-4ubuntu1 Target: x86_64-pc-linux-gnu Thread model: posix InstalledDir: /usr/bin |
Summary
This PR extracts the key value and gradient calculations from the
std_normal_lcdfheader into reusable functions that other normal-cdf-related functions can delegate to, removing a lot of duplicated/inconsistent implementations.The value and gradient calculations are also updated to replace the taylor expansions with the Cody rational approximations (previously only used for the tail).
Two other bugs were found & fixed for
var_value<Eigen::VectorXd>handling:as_array_or_scalar()now returns an array unchanged instead of re-wrapping itref_type_ifkeeps an rvaluearena_matrixby value rather than referencing a dead temporaryThe valid ranges of inputs (before over/underflow) across all functions are expanded:
While performance is either unchanged or significantly improved across inputs:
Full timing comparisons
double (values only)
var_value (Eigen::Matrix<var_value> y, var_value parameters, gradient)
var_valueEigen::Matrix (var_valueEigen::VectorXd y, var_value parameters, gradient)
fvar (tangent)
Tests
Additional tests exercising the increased numerical stability are added, as well new
mixtests for the distributionsSide Effects
N/A
Release notes
Increased numerical stability and performance of the
normal,std_normal,lognormalandexp_mod_normal(LC)CDF functionsChecklist
Copyright holder: (fill in copyright holder information)
The copyright holder is typically you or your assignee, such as a university or company. By submitting this pull request, the copyright holder is agreeing to the license the submitted work under the following licenses:
- Code: BSD 3-clause (https://opensource.org/licenses/BSD-3-Clause)
- Documentation: CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
the basic tests are passing
./runTests.py test/unit)make test-headers)make test-math-dependencies)make doxygen)make cpplint)the code is written in idiomatic C++ and changes are documented in the doxygen
the new changes are tested