Tell HN: Litellm 1.82.7 and 1.82.8 on PyPI are compromised
TL;DR Highlight
Malicious .pth files stealing credentials were inserted into LiteLLM PyPI packages versions 1.82.7 and 1.82.8. A supply chain attack that auto-executes on Python interpreter startup — without any import — giving it a wide blast radius.
Who Should Read
Backend developers and MLOps engineers using LiteLLM — specifically anyone who pip-installed litellm in an AI service development environment. Immediate action required.
Core Mechanics
- .pth files in Python's site-packages directory are automatically executed when the Python interpreter starts — no import statement needed. This makes them an unusually dangerous vector for supply chain attacks.
- The malicious code in versions 1.82.7 and 1.82.8 collected environment variables (including API keys, cloud credentials, and database URLs) and exfiltrated them to an external server.
- Any environment where litellm was installed — including Docker containers, virtual environments, and CI/CD pipelines — may have had credentials exfiltrated at every Python process startup.
- The attack was discovered and the malicious versions were yanked from PyPI, but anyone who installed those specific versions between release and yanking is affected.
- Mitigation: immediately upgrade to a clean version, rotate all credentials accessible in environments where 1.82.7 or 1.82.8 was installed.
Evidence
- The security researcher who discovered the attack shared the decompiled malicious code, confirming the .pth execution mechanism and the exfiltration endpoint.
- LiteLLM maintainers published an incident response within hours, confirming the attack, which versions were affected, and recommending immediate upgrade and credential rotation.
- The attack hit particularly hard in AI development environments where litellm is used as a routing layer — these environments typically have credentials for many AI providers (OpenAI, Anthropic, etc.) in their environment variables.
- Commenters raised the broader point: LiteLLM is exactly the kind of high-value target for supply chain attacks — widely used in AI infrastructure, often installed with broad permissions.
How to Apply
- Immediately: pip install --upgrade litellm to get a clean version. Then rotate all credentials that were accessible as environment variables in affected environments.
- Check your pip history or requirements.txt locks to determine if you installed 1.82.7 or 1.82.8. If unclear, treat it as compromised and rotate anyway.
- Add a .pth file scanner to your dependency audit process — tools like pip-audit and safety don't currently detect malicious .pth files, so consider adding a custom check.
- For AI service environments, store sensitive credentials in a secrets manager (AWS Secrets Manager, Vault) rather than environment variables — this limits blast radius from env var exfiltration attacks.
Code Example
# Script to check for malicious package
pip download litellm==1.82.8 --no-deps -d /tmp/check
python3 -c "
import zipfile, os
whl = '/tmp/check/' + [f for f in os.listdir('/tmp/check') if f.endswith('.whl')][0]
with zipfile.ZipFile(whl) as z:
pth = [n for n in z.namelist() if n.endswith('.pth')]
print('PTH files:', pth) # Should be an empty list if clean
for p in pth:
print(z.read(p)[:300]) # Inspect contents
"
# Pin to a safe version
pip install litellm==1.82.6
# Pin version in requirements.txt
# litellm==1.82.6Terminology
Related Papers
Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
동일한 모델과 작업 환경에서 Claude Code와 OpenCode의 실제 토큰 사용량을 API 레벨에서 측정한 결과, Claude Code가 시스템 프롬프트 오버헤드만으로 OpenCode 대비 4.7배 더 많은 토큰을 소비한다는 것을 확인했다.
Mesh LLM: distributed AI computing on iroh
사무실, 집, 클라우드에 흩어진 GPU들을 하나의 OpenAI 호환 API로 묶어주는 분산 LLM 실행 시스템으로, 비싼 API 비용 없이 큰 모델을 직접 운영할 수 있다.
Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)
내 LLM API 비용이 어디서 새는지 로컬에서 분석해주는 오픈소스 CLI 도구로, 비싼 모델 대신 저렴한 모델로 전환 가능한 호출을 골라낸다.
Jamesob's guide to running SOTA LLMs locally
2천 달러짜리 RTX 3090 한 장부터 4만 달러짜리 RTX PRO 6000 4장 셋업까지, 로컬에서 최신 LLM을 직접 돌리는 방법을 하드웨어 선택·구성·실행 설정까지 통째로 정리한 실전 가이드다.
Faster embeddings: how we rebuilt the ONNX path in Manticore
Manticore Search가 기존 SentenceTransformers/Candle 백엔드를 ONNX Runtime으로 교체해 텍스트 임베딩 생성 속도를 평균 14배 향상시켰다. 별도 모델 서비스 없이 DB 내부에서 직접 임베딩을 처리하는 구조에서 INSERT 속도가 곧 임베딩 속도이기 때문에 이 개선은 실질적인 ingest 처리량 향상으로 직결된다.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
멀티벡터 검색 모델의 문서 벡터를 1비트 이진값으로 압축하고 쿼리 벡터만 int8로 유지하는 비대칭 양자화 기법으로, 스토리지를 97% 줄이면서 검색 품질 손실을 0.61점(NDCG@10 기준)에 그치게 만든 실제 프로덕션 적용 사례다.