Faster asin() was hiding in plain sight
TL;DR Highlight
While optimizing a ray tracer, the author hand-implemented asin() using Taylor series and Padé approximation — faster and more accurate than the system library version.
Who Should Read
Graphics programmers, game engine developers, and performance engineers who care about floating-point performance and numerical methods.
Core Mechanics
- For ray tracing performance, the author needed a fast asin() (arcsine) function and found the system math library version was slower than necessary for their precision requirements.
- Taylor series approximation: expanding asin(x) around x=0 gives a polynomial that's fast to compute but loses accuracy near x=±1.
- Padé approximation: a rational function (polynomial/polynomial) that provides better accuracy than Taylor series for the same number of terms, especially near the boundaries of the domain.
- The custom implementation was faster than the system asin() while meeting the ray tracer's precision needs — the key insight being that ray tracers often don't need full IEEE 754 double precision.
- This kind of micro-optimization matters when asin() is called millions of times per frame in tight inner loops.
Evidence
- The author shared benchmark results comparing their Padé approximation against std::asin() on multiple platforms, showing 2-3x speedup.
- HN commenters with numerical methods background discussed the tradeoffs between Taylor, Chebyshev, and Padé approximations for different domains.
- Some noted that SIMD-vectorized versions of these approximations (using AVX2/NEON) can push even larger speedups.
- Others pointed out that modern CPUs have fast hardware asin() in the FPU, and the wins depend heavily on the CPU microarchitecture.
How to Apply
- If you're calling math functions millions of times per frame, profile to confirm they're actual bottlenecks before optimizing — the standard library is often fast enough.
- When you do need a custom approximation, Padé approximations generally outperform Taylor series for the same computation cost when you need accuracy across a wider domain.
- Consider the precision requirements carefully: graphics often tolerates 1e-6 relative error, physics simulation might need 1e-10 — match the approximation to the requirement.
- Check if SIMD-vectorized math libraries (Intel SVML, Sleef) already provide what you need before writing your own approximation.
Code Example
// NVIDIA CG library-based fast_asin (Stack Overflow: https://stackoverflow.com/a/26030435)
// Original source: Hastings 1955 / Abramowitz & Stegun 4.4.45
float fast_asin(float x) {
float negate = float(x < 0);
x = abs(x);
float ret = -0.0187293f;
ret *= x;
ret += 0.0742610f;
ret *= x;
ret -= 0.2121144f;
ret *= x;
ret += 1.5707288f;
ret = 3.14159265358979f * 0.5f - sqrt(1.0f - x) * ret;
return ret - 2 * negate * ret;
}
// Taylor series-based 4th-order approximation (author's own implementation, ~5% improvement, valid only in range -0.8 ~ 0.8)
double _asin_approx_private(const double x) {
if ((x < -0.8) || (x > 0.8)) {
return std::asin(x); // fallback
}
constexpr double a = 0.5;
constexpr double b = a * 0.75;
constexpr double c = b * (5.0 / 6.0);
constexpr double d = c * (7.0 / 8.0);
const double aa = (x * x * x) / 3.0;
const double bb = (x * x * x * x * x) / 5.0;
const double cc = (x * x * x * x * x * x * x) / 7.0;
const double dd = (x * x * x * x * x * x * x * x * x) / 9.0;
return x + (a * aa) + (b * bb) + (c * cc) + (d * dd);
}Terminology
Related Papers
Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
동일한 모델과 작업 환경에서 Claude Code와 OpenCode의 실제 토큰 사용량을 API 레벨에서 측정한 결과, Claude Code가 시스템 프롬프트 오버헤드만으로 OpenCode 대비 4.7배 더 많은 토큰을 소비한다는 것을 확인했다.
Mesh LLM: distributed AI computing on iroh
사무실, 집, 클라우드에 흩어진 GPU들을 하나의 OpenAI 호환 API로 묶어주는 분산 LLM 실행 시스템으로, 비싼 API 비용 없이 큰 모델을 직접 운영할 수 있다.
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
내 LLM API 비용이 어디서 새는지 로컬에서 분석해주는 오픈소스 CLI 도구로, 비싼 모델 대신 저렴한 모델로 전환 가능한 호출을 골라낸다.
Jamesob's guide to running SOTA LLMs locally
2천 달러짜리 RTX 3090 한 장부터 4만 달러짜리 RTX PRO 6000 4장 셋업까지, 로컬에서 최신 LLM을 직접 돌리는 방법을 하드웨어 선택·구성·실행 설정까지 통째로 정리한 실전 가이드다.
Faster embeddings: how we rebuilt the ONNX path in Manticore
Manticore Search가 기존 SentenceTransformers/Candle 백엔드를 ONNX Runtime으로 교체해 텍스트 임베딩 생성 속도를 평균 14배 향상시켰다. 별도 모델 서비스 없이 DB 내부에서 직접 임베딩을 처리하는 구조에서 INSERT 속도가 곧 임베딩 속도이기 때문에 이 개선은 실질적인 ingest 처리량 향상으로 직결된다.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
멀티벡터 검색 모델의 문서 벡터를 1비트 이진값으로 압축하고 쿼리 벡터만 int8로 유지하는 비대칭 양자화 기법으로, 스토리지를 97% 줄이면서 검색 품질 손실을 0.61점(NDCG@10 기준)에 그치게 만든 실제 프로덕션 적용 사례다.
Related Resources
- Original: Faster asin() Was Hiding In Plain Sight
- Stack Overflow: Fast asin approximation (NVIDIA CG-based)
- Introduction to Chebyshev Approximation (embeddedrelated.com)
- Remez Algorithm (Wikipedia)
- Robin Green: Faster Math Functions (GDC slide deck 1)
- iquilezles: Avoiding Trigonometry (noacos)
- glibc asin implementation source
- Intel AVX-512 asin intrinsic guide