AI网关LiteLLM那些事儿
近期我们上线了LiteLLM,社区活跃,star最高。主要解决,1简化接入 2计费限流。天有不测风云,上线后被外部脱裤!使用https://www.sysdig.com/blog/cve-2026-42208-targeted-sql-injection-against-litellms-authentication-path-discovered-36-hours-following-vulnerability-disclosure 绕过鉴权,数据库中查看到上游key被拿到,一顿轮转。加白升级版本。
随着使用过程的深入,架构也由单台-多台+数据库-redis,配置也有不少讲究
当前架构
┌───────────┐
│ ALB │
│ (公网入口) │
└─────┬─────┘
╱ ╲
┌────────▼───┐ ┌────▼────────┐
│ gateway01 │ │ gateway02 │
│ 192.168. │ │ 192.168. │
│ 1.126 │ │ 3.108 │
│ Gunicorn │ │ Gunicorn │
│ 4 workers │ │ 4 workers │
└──────┬─────┘ └──────┬──────┘
│ cron 每分钟 │
│ config 同步 │
│ ◄──────────── │
└────────┬───────┘
┌─────▼──────┐
│ Redis 7.0 │
│ 阿里云 │
└─────┬──────┘
┌─────▼──────┐
│ PostgreSQL │
│ 阿里云 RDS │
└────────────┘各组件职责
| 组件 | 说明 |
|---|---|
| ALB | 阿里云负载均衡,将流量分发到两台实例 |
| aigateway01 | 主配置节点,config.yaml 在此维护 |
| aigateway02 | 从节点,通过 cron 每分钟从 01 同步 config.yaml |
| Redis 7.0 | 共享缓存、限流计数器、模型故障冷却状态 |
| PostgreSQL | 存储用户/Key/消费记录等持久化数据 |
配置文件详情
4.1 config.yaml
路径:/opt/litellm/config.yaml
说明:LiteLLM 核心配置,01 为主,02 通过 cron 同步
# ===== 通用设置 =====
general_settings:
database_url: os.environ/DATABASE_URL
master_key: os.environ/LITELLM_MASTER_KEY
store_model_in_db: true
# 不在 spend_logs 中存储 prompt 内容,降低内存和存储消耗
store_prompts_in_spend_logs: false
# 批量写入消费记录到 DB,每60秒一次,减轻 DB 压力(官方生产建议)
proxy_batch_write_at: 60
# 连接池限制:MAX_DB_CONNECTIONS / (实例数 x workers) = 约5
database_connection_pool_limit: 5
# VPC 内网部署,DB 故障时仍允许请求通过(官方生产建议)
allow_requests_on_db_unavailable: true
# ===== LiteLLM 核心设置 =====
litellm_settings:
# 启用 Redis 缓存(通过环境变量 REDIS_HOST/PORT/PASSWORD 自动连接)
cache: true
cache_params:
type: redis
# 请求超时 600 秒(默认 6000 秒太长,上游挂掉会卡很久)
request_timeout: 600
# 生产环境关闭 debug 日志
set_verbose: false
# ===== 模型列表(通过 DB 管理,此处留空) =====
model_list: []
# ===== 路由设置 =====
router_settings:
# 连续失败 2 次后冷却该模型
allowed_fails: 2
# 冷却时间 30 秒
cooldown_time: 30
# 模型别名映射
model_group_alias:
gpt-5.4-mini: openai/gpt-5.4-mini
gpt-5.4-nano: openai/gpt-5.4-nano
gpt-5.4: openai/gpt-5.4-fb
openai/gpt-5.4: openai/gpt-5.4-fb
# 以下状态码触发自动重试
retry_on_status_codes:
- 401
- 403
- 408
- 429
- 500
- 502
- 503
- 504config.yaml 逐项说明
| 配置项 | 值 | 本次新增 | 说明 |
|---|---|---|---|
database_url | 环境变量引用 | 从环境变量读取 PG 连接串 | |
master_key | 环境变量引用 | API 主密钥 | |
store_model_in_db | true | 模型配置存 DB,支持 UI 动态管理 | |
store_prompts_in_spend_logs | false | 不存 prompt 原文,节省存储 | |
proxy_batch_write_at | 60 | ✅ | 消费记录每 60 秒批量写入 DB,原来每次请求都写 |
database_connection_pool_limit | 5 | ✅ | 每个 worker 最多 5 个 DB 连接。2 实例 × 4 worker × 5 = 最多 40 连接 |
allow_requests_on_db_unavailable | true | ✅ | DB 故障时不拒绝请求,适用于 VPC 内网部署 |
cache | true | ✅ | 启用 Redis 缓存 |
cache_params.type | redis | ✅ | 缓存后端为 Redis(连接信息由环境变量自动注入) |
request_timeout | 600 | ✅ | 请求超时 10 分钟(默认 6000 秒=100 分钟,太长) |
set_verbose | false | ✅ | 关闭 debug 日志 |
allowed_fails | 2 | 模型连续失败 2 次后进入冷却 | |
cooldown_time | 30 | 冷却时间 30 秒 | |
model_group_alias | 见上 | 模型名称别名映射 | |
retry_on_status_codes | 401-504 | 这些状态码会触发自动重试 |
4.2 docker-compose.yml
路径:/opt/litellm/docker-compose.yml
说明:两台实例配置相同
# ===== LiteLLM AI Gateway 生产配置 =====
services:
litellm:
image: docker.litellm.ai/berriai/litellm:main-v1.82.3-stable
container_name: litellm
restart: always
deploy:
resources:
limits:
memory: 4g
reservations:
memory: 512m
volumes:
- ./config.yaml:/app/config.yaml
command:
- --config=/app/config.yaml
- --num_workers
- "4"
# 使用 Gunicorn 替代默认 uvicorn,worker 回收更稳定(官方生产建议)
- "--run_gunicorn"
# 每处理 10000 个请求后回收 worker,防止内存泄漏
- --max_requests_before_restart
- "10000"
ports:
- 0.0.0.0:4000:4000
env_file:
- .env
environment:
DATABASE_URL: postgresql://litellm:****@pgm-****.pgsql.singapore.rds.aliyuncs.com:5432/aigateway?connection_limit=5&pool_timeout=10
STORE_MODEL_IN_DB: 'True'
LITELLM_MODE: 'PRODUCTION'
# 减少 FastAPI 默认 INFO 日志输出(官方生产建议)
LITELLM_LOG: 'ERROR'
healthcheck:
test:
- CMD-SHELL
- python3 -c "import urllib.request; urllib.request.urlopen('http://localhost:4000/health/liveliness')"
interval: 30s
timeout: 10s
retries: 3
start_period: 40s
networks:
- litellm-net
logging:
driver: json-file
options:
max-size: 500m
max-file: '3'
# ===== Prometheus 监控 =====
prometheus:
image: prom/prometheus
container_name: litellm_prometheus
restart: always
volumes:
- prometheus_data:/prometheus
- ./prometheus.yml:/etc/prometheus/prometheus.yml
ports:
- 127.0.0.1:9090:9090
command:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.path=/prometheus
- --storage.tsdb.retention.time=15d
networks:
- litellm-net
logging:
driver: json-file
options:
max-size: 500m
max-file: '3'
volumes:
prometheus_data:
name: litellm_prometheus_data
networks:
litellm-net:
driver: bridgedocker-compose.yml 本次变更项
| 配置项 | 值 | 说明 |
|---|---|---|
--run_gunicorn | 新增 | 使用 Gunicorn 管理 worker 进程,回收机制比 uvicorn 更成熟稳定 |
LITELLM_LOG | ERROR | 关闭 FastAPI 默认的 INFO 级别请求日志,减少大量无用输出 |
4.3 .env
路径:/opt/litellm/.env
说明:两台实例配置相同,敏感信息通过环境变量注入
LITELLM_MASTER_KEY=sk-litellm-****
DATABASE_URL=postgresql://litellm:****@pgm-****.pgsql.singapore.rds.aliyuncs.com:5432/aigateway
STORE_MODEL_IN_DB=True
LITELLM_LOG=INFO
UI_USERNAME=admin
UI_PASSWORD=****
OPENAI_API_KEY=sk-proj-****
# Redis 连接(分开写法,比 REDIS_URL 快 ~80 RPS)
REDIS_HOST=aigateway.xxxxx.redis.sunmi.com
REDIS_PORT=6379
REDIS_USERNAME=aigateway
REDIS_PASSWORD='VbRyLVl)Ixxxxxxx'[!WARNING]
Redis 密码包含特殊字符)、&、(,在.env中必须用单引号包裹,否则会被 shell 解析导致连接失败。
“这是好事”
底裤都被看透;(