近期我们上线了LiteLLM,社区活跃,star最高。主要解决,1简化接入 2计费限流。天有不测风云,上线后被外部脱裤!使用https://www.sysdig.com/blog/cve-2026-42208-targeted-sql-injection-against-litellms-authentication-path-discovered-36-hours-following-vulnerability-disclosure 绕过鉴权,数据库中查看到上游key被拿到,一顿轮转。加白升级版本。
随着使用过程的深入,架构也由单台-多台+数据库-redis,配置也有不少讲究

当前架构

                       ┌───────────┐
                       │    ALB    │
                       │ (公网入口) │
                       └─────┬─────┘
                      ╱             ╲
            ┌────────▼───┐    ┌────▼────────┐
            │ gateway01  │    │ gateway02   │
            │ 192.168.   │    │ 192.168.    │
            │ 1.126      │    │ 3.108       │
            │ Gunicorn   │    │ Gunicorn    │
            │ 4 workers  │    │ 4 workers   │
            └──────┬─────┘    └──────┬──────┘
                   │   cron 每分钟    │
                   │   config 同步    │
                   │  ◄──────────── │
                   └────────┬───────┘
                      ┌─────▼──────┐
                      │  Redis 7.0 │
                      │  阿里云     │
                      └─────┬──────┘
                      ┌─────▼──────┐
                      │ PostgreSQL │
                      │  阿里云 RDS │
                      └────────────┘

各组件职责

组件说明
ALB阿里云负载均衡,将流量分发到两台实例
aigateway01主配置节点,config.yaml 在此维护
aigateway02从节点,通过 cron 每分钟从 01 同步 config.yaml
Redis 7.0共享缓存、限流计数器、模型故障冷却状态
PostgreSQL存储用户/Key/消费记录等持久化数据

配置文件详情

4.1 config.yaml

路径:/opt/litellm/config.yaml
说明:LiteLLM 核心配置,01 为主,02 通过 cron 同步
# ===== 通用设置 =====
general_settings:
  database_url: os.environ/DATABASE_URL
  master_key: os.environ/LITELLM_MASTER_KEY
  store_model_in_db: true
  # 不在 spend_logs 中存储 prompt 内容,降低内存和存储消耗
  store_prompts_in_spend_logs: false
  # 批量写入消费记录到 DB,每60秒一次,减轻 DB 压力(官方生产建议)
  proxy_batch_write_at: 60
  # 连接池限制:MAX_DB_CONNECTIONS / (实例数 x workers) = 约5
  database_connection_pool_limit: 5
  # VPC 内网部署,DB 故障时仍允许请求通过(官方生产建议)
  allow_requests_on_db_unavailable: true

# ===== LiteLLM 核心设置 =====
litellm_settings:
  # 启用 Redis 缓存(通过环境变量 REDIS_HOST/PORT/PASSWORD 自动连接)
  cache: true
  cache_params:
    type: redis
  # 请求超时 600 秒(默认 6000 秒太长,上游挂掉会卡很久)
  request_timeout: 600
  # 生产环境关闭 debug 日志
  set_verbose: false

# ===== 模型列表(通过 DB 管理,此处留空) =====
model_list: []

# ===== 路由设置 =====
router_settings:
  # 连续失败 2 次后冷却该模型
  allowed_fails: 2
  # 冷却时间 30 秒
  cooldown_time: 30
  # 模型别名映射
  model_group_alias:
    gpt-5.4-mini: openai/gpt-5.4-mini
    gpt-5.4-nano: openai/gpt-5.4-nano
    gpt-5.4: openai/gpt-5.4-fb
    openai/gpt-5.4: openai/gpt-5.4-fb
  # 以下状态码触发自动重试
  retry_on_status_codes:
  - 401
  - 403
  - 408
  - 429
  - 500
  - 502
  - 503
  - 504

config.yaml 逐项说明

配置项值本次新增说明
database_url环境变量引用 从环境变量读取 PG 连接串
master_key环境变量引用 API 主密钥
store_model_in_dbtrue 模型配置存 DB,支持 UI 动态管理
store_prompts_in_spend_logsfalse 不存 prompt 原文,节省存储
proxy_batch_write_at60✅消费记录每 60 秒批量写入 DB,原来每次请求都写
database_connection_pool_limit5✅每个 worker 最多 5 个 DB 连接。2 实例 × 4 worker × 5 = 最多 40 连接
allow_requests_on_db_unavailabletrue✅DB 故障时不拒绝请求,适用于 VPC 内网部署
cachetrue✅启用 Redis 缓存
cache_params.typeredis✅缓存后端为 Redis(连接信息由环境变量自动注入)
request_timeout600✅请求超时 10 分钟(默认 6000 秒=100 分钟,太长)
set_verbosefalse✅关闭 debug 日志
allowed_fails2 模型连续失败 2 次后进入冷却
cooldown_time30 冷却时间 30 秒
model_group_alias见上 模型名称别名映射
retry_on_status_codes401-504 这些状态码会触发自动重试

4.2 docker-compose.yml

路径:/opt/litellm/docker-compose.yml
说明:两台实例配置相同
# ===== LiteLLM AI Gateway 生产配置 =====
services:
  litellm:
    image: docker.litellm.ai/berriai/litellm:main-v1.82.3-stable
    container_name: litellm
    restart: always
    deploy:
      resources:
        limits:
          memory: 4g
        reservations:
          memory: 512m
    volumes:
    - ./config.yaml:/app/config.yaml
    command:
    - --config=/app/config.yaml
    - --num_workers
    - "4"
    # 使用 Gunicorn 替代默认 uvicorn,worker 回收更稳定(官方生产建议)
    - "--run_gunicorn"
    # 每处理 10000 个请求后回收 worker,防止内存泄漏
    - --max_requests_before_restart
    - "10000"
    ports:
    - 0.0.0.0:4000:4000
    env_file:
    - .env
    environment:
      DATABASE_URL: postgresql://litellm:****@pgm-****.pgsql.singapore.rds.aliyuncs.com:5432/aigateway?connection_limit=5&pool_timeout=10
      STORE_MODEL_IN_DB: 'True'
      LITELLM_MODE: 'PRODUCTION'
      # 减少 FastAPI 默认 INFO 日志输出(官方生产建议)
      LITELLM_LOG: 'ERROR'
    healthcheck:
      test:
      - CMD-SHELL
      - python3 -c "import urllib.request; urllib.request.urlopen('http://localhost:4000/health/liveliness')"
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 40s
    networks:
    - litellm-net
    logging:
      driver: json-file
      options:
        max-size: 500m
        max-file: '3'
  # ===== Prometheus 监控 =====
  prometheus:
    image: prom/prometheus
    container_name: litellm_prometheus
    restart: always
    volumes:
    - prometheus_data:/prometheus
    - ./prometheus.yml:/etc/prometheus/prometheus.yml
    ports:
    - 127.0.0.1:9090:9090
    command:
    - --config.file=/etc/prometheus/prometheus.yml
    - --storage.tsdb.path=/prometheus
    - --storage.tsdb.retention.time=15d
    networks:
    - litellm-net
    logging:
      driver: json-file
      options:
        max-size: 500m
        max-file: '3'
volumes:
  prometheus_data:
    name: litellm_prometheus_data
networks:
  litellm-net:
    driver: bridge

docker-compose.yml 本次变更项

配置项值说明
--run_gunicorn新增使用 Gunicorn 管理 worker 进程,回收机制比 uvicorn 更成熟稳定
LITELLM_LOGERROR关闭 FastAPI 默认的 INFO 级别请求日志,减少大量无用输出

4.3 .env

路径:/opt/litellm/.env
说明:两台实例配置相同,敏感信息通过环境变量注入
LITELLM_MASTER_KEY=sk-litellm-****
DATABASE_URL=postgresql://litellm:****@pgm-****.pgsql.singapore.rds.aliyuncs.com:5432/aigateway
STORE_MODEL_IN_DB=True
LITELLM_LOG=INFO
UI_USERNAME=admin
UI_PASSWORD=****
OPENAI_API_KEY=sk-proj-****

# Redis 连接(分开写法,比 REDIS_URL 快 ~80 RPS)
REDIS_HOST=aigateway.xxxxx.redis.sunmi.com
REDIS_PORT=6379
REDIS_USERNAME=aigateway
REDIS_PASSWORD='VbRyLVl)Ixxxxxxx'
[!WARNING]
Redis 密码包含特殊字符 )、&、(,在 .env 中必须用单引号包裹,否则会被 shell 解析导致连接失败。

标签: aigateway, litellm, 安全, 高可用

已有 2 条评论

  1. ‌“这是好事”

    1. 底裤都被看透;(

添加新评论