现网(2026-10-06):TP=2 · TQ3 KV · 131072 · 9 副本 · 调用说明见 27b 用法(含并发/输出表)。运维仓库 llm-27b-up。

2026-09-10 · L1 + L2 · 已回滚

27b 服务端升级记录 —— 一次被推翻的升级

这是失败记录,不是发布说明。L1(util 0.91→0.944)+ L2(max-model-len 131072→262144)在 2026-09-10 19:45 推到生产, 19:50 被第一个真实长请求(prompt_tokens=93199)打崩,19:55 字节级回滚, 20:02 用同一请求复验通过。生产当前仍是 0.91 / 131072 / 单请求 128K。 下面所有「改后」列的 256K 数字都是已回滚、不再生效的历史值。

改前(保留,即现值)改后(已回滚)
--gpu-memory-utilization0.910.944
--max-model-len131072262144
单请求上下文128K256K
启动期 KV 池(生产实测)679,191 token855,417 token(+25.9%)
运行期 93K prefillPASS(answer=7Q3Z)CUDA OOM → 容器重启
结论生产用它拒绝

最贵的一条经验:启动日志那三行(Available KV cache memory / GPU KV cache size / Maximum concurrency)全绿也不能证明配置可用。 那三行只算「权重 + KV 池 + cudagraph + 非 torch 开销」,一行都没算运行期 prefill 的激活峰值。

为什么当时以为是安全的

util 0.91→0.944 不是我们拍的,是 vLLM 启动日志自己算出来的建议值:

gpu_worker:553  util 0.9100 ≡ 0.8760 (cudagraph profiling on);
                「要维持改之前同样的 KV,请设到 0.9440」

0.91 只等于「关掉 cudagraph profiling 时的 0.876」,那 0.034 是 profiler 预留的额度。 当时把它读成「引擎自己认为安全的额度,可以要回来」。

max-model-len 131072→262144 不花 1 字节显存:模型原生 max_position_embeddings=262144、rope_type="default"(不需要 yarn/linear 缩放)。 131072 是我们自己设的上限。

这两条推断单独看都对,合起来的结论错了:0.944 要回来的那 0.67 GiB 并不是「多余额度」, 而是 prefill 激活峰值的余量;而 262144 把这个余量能承受的 prefill 长度上限抬高了。

执行与实测

Canary 只验了启动日志

canary vllm-qwen38-27b-tq4(1 副本)改前是 32768 / 0.80 / 9 seqs / 8192 batched, 不是生产的同构克隆 —— 它只验证「这两个 flag 能不能在这个模型上安全落地」。 新 pod 实测(19:35:11):

Checkpoint size: 16.19 GiB
Model loading took 8.07 GiB memory and 165.696784 seconds
Estimated CUDA graph memory: 0.44 GiB total          (该 pod capture sizes 只到 8)
Available KV cache memory: 9.01 GiB
GPU KV cache size: 965,793 tokens
Maximum concurrency for 262,144 tokens per request: 3.68x
Application startup complete.
(无 CUDA out of memory / 无 graph capture 报错)

canary 起来了、三行全绿,于是被判定为「通过」。但运行期门禁 longprefill_probe.py 没跑(当时它还不存在)。这是事故的第一层原因。

生产跳过门禁直接推

第一个新 pod vllm-qwen38-27b-lora-754758c4ff-npqv4(19:45:11):

Checkpoint size: 16.19 GiB
Model loading took 8.07 GiB memory and 177.651499 seconds
Padding mamba page size by 167.69% ...
Setting attention block size to 8192 tokens ...
Estimated CUDA graph memory: 0.67 GiB total
Available KV cache memory: 7.98 GiB
GPU KV cache size: 855,417 tokens
Maximum concurrency for 262,144 tokens per request: 3.26x
CUDA graph pool memory: 0.61 GiB (actual), 0.67 GiB (estimated), difference: 0.05 GiB (8.6%)
Application startup complete.
(无 CUDA out of memory / 无 graph capture 报错)

改前/改后对照(同一 Deployment、同一份 args,只差那两个 flag):

改前(0.91 / 131072)改后(0.944 / 262144)Δ
Available KV cache memory7.31 GiB7.98 GiB+0.67 GiB
GPU KV cache size679,191 token855,417 token+176,226(+25.9%)
满长并发 @128K5.18x6.53x+26.0%
满长并发 @256K—(上限 128K)3.26x—
num_gpu_blocks114124+8.8%
token/GiB92,913107,196+15.4%
CUDA graph 池0.67 估 / 0.61 实0.67 估 / 0.61 实0

数字取自两个 pod 的 /metrics 的 vllm:cache_config_info; 两边 block_size=16、mamba_block_size=16、mamba_cache_mode=align、 cache_dtype=turboquant_4bit_nc、kv_cache_dtype_skip_layers 完全一致。

崩溃:19:50:45,第一个真实长请求

对 npqv4 发 prompt_tokens=93199 → HTTP 500:

(Worker_TP0 pid=577) ERROR 09-10 19:50:45 [multiproc_executor.py:1004]
  File ".../vllm/model_executor/layers/fla/ops/chunk_delta_h.py", line 357,
       in chunk_gated_delta_rule_fwd_h
    v_new = torch.empty_like(u) if save_new_value else None
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 96.00 MiB.
GPU 0 has a total capacity of 19.58 GiB of which 94.88 MiB is free.
Including non-PyTorch memory, this process has 19.48 GiB memory in use.

同一个请求打到旧配置 pod 755678f46c-6sbxc(0.91 / 131072): prompt_tokens=93199 finish=stop answer='7Q3Z' —— PASS。

把每卡显存账按两个 util 各算一遍(3080,分母 20480 MiB,可用 19.58 GiB):

项0.91(改前 / 现值)0.944(已回滚)
预算 = util × 20480 MiB18.20 GiB18.88 GiB
权重8.078.07
CUDA graph 池(实测)0.610.61
KV 池(启动时预分配)7.317.98
非 torch 开销(CUDA context / NCCL / cuBLAS / LoRA)2.212.22
合计18.20 ✓18.88 ✓
留给运行期激活的余量1.38 GiB0.70 GiB

那 0.67 GiB 不是「白捡的额度」,它就是 prefill 激活峰值的安全垫。启动阶段的账一行都没算运行期 (chunked prefill 的中间张量、GDN 的 u/v_new、attention workspace)。 崩溃点恰好卡在临界值上:只剩 94.88 MiB,而 empty_like(u) 要 96.00 MiB。

单变量 A/B:肇事者是 L2,不是 util

因为 93199 < 131072,max-model-len 本身没有拒绝这个请求,所以第一嫌疑是 util。 把 canary 对齐成生产的同构克隆(48 层 skip-layers、--max-num-seqs 64、 --max-num-batched-tokens 16384、--cudagraph-capture-sizes 1 2 4 8 16 32), 逐 flag 比对只差 --gpu-memory-utilization 与 --served-model-name, 每次只留一个变量,跑同一道门禁 longprefill_probe.py <pod> 60000 80000:

实验utilmax-model-len其余门禁结果
A0.944131072与生产一致PASS(93,199 与 124,637 两步全过,restartCount 0)
B0.91262144与生产一致FAIL(93,199 一步 HTTP_500,restartCount 1→2)

实验 A(util 0.944 / max-model-len 131072)canary ...-fz5g6 启动实测:

Available KV cache memory: 7.98 GiB
GPU KV cache size: 738,769 tokens
Maximum concurrency for 131,072 tokens per request: 5.64x
=== target=60000 ip=10.244.209.45 chars=324643 restarts_ready=0:true 20:15:00
OK prompt_tokens=93199 completion=5 finish=stop elapsed=113s answer='7Q3Z' PASS
    after: restarts_ready=0:true 20:16:53
=== target=80000 ip=10.244.209.45 chars=433247 restarts_ready=0:true 20:16:53
OK prompt_tokens=124637 completion=5 finish=stop elapsed=129s answer='7Q3Z' PASS
    after: restarts_ready=0:true 20:19:03
SUMMARY [[60000, "PASS"], [80000, "PASS"]]
DONE 20:19:03

这一步否掉了第一嫌疑:同一个 93K 请求、同样的 --max-model-len 131072,util 抬到 0.944 之后照样过,restartCount 仍是 0,且 Available KV cache memory 也是 7.98 GiB (和生产崩掉那版一模一样)。「0.944 吃掉了 prefill 安全垫」这个解释单独站不住。 顺带修正口径:738,769 ÷ 7.98 = 92,578 token/GiB,回到基线 92,913 附近。

实验 B(util 0.91 / max-model-len 262144)canary ...-4bsnt 部署 args 只改了 --max-model-len,util 保持 0.91。 启动实测(20:39:04–20:39:45):

Available KV cache memory: 7.31 GiB
GPU KV cache size: 786,432 tokens
Maximum concurrency for 262,144 tokens per request: 3.00x
...
INFO: Application startup complete.            # 20:39:45

启动是绿的,786,432 ÷ 7.31 = 107,582 token/GiB,与生产崩掉那版基本一致。 同一道门禁:

=== target=60000 ip=10.244.218.132 chars=324643 restarts_ready=1:true 20:39:xx
93199 HTTP_500 elapsed=301s
    after: restarts_ready=2:false
SUMMARY [[60000, "HTTP_500"]]

第一步 93,199 token 就挂了。日志(kubectl logs --previous)里的崩溃链 不是 OOM,也不是 CUDA 错:

shm_broadcast.py:705        ...  ×3(20:42 / 20:43 / 20:44,每次 60s 超时)
core.py:1233                TimeoutError: RPC call to sample_tokens timed out   # 20:44:46
async_llm.py:704            EngineDeadError

即引擎在 93K prefill 上先卡死(worker 与 scheduler 之间的 sample_tokens RPC 连续三次 60s 无响应),scheduler 把引擎判死后整进程退出。控制组(实验 A 与生产 nx7mc) 在同一请求上是 113–126s 正常返回答案。

结论:--max-model-len 262144 单独就足以让 93K prefill 崩, 与 util 无关;生产那次的 OOM 主因是 L2,L1 只是把安全垫从 1.38 GiB 压到 0.70 GiB、放大了后果。

另记一笔(实验 B 首轮,非本轮门禁数据):第一次起 exp B 时崩在 turboquant_attn.py:834 _continuation_prefill 的 k_flat = k_flat @ Pi_half → CUBLAS_STATUS_EXECUTION_FAILED(源 /tmp/expB.log)。重跑后崩点变成上面的 RPC 超时, 说明 262,144 这条路径在 prefill 阶段还有非确定性的失败面(两次崩点不同),不是单一 bug。

一条被推翻两次的口径

同是 TQ4 KV,几个部署量出来的 token/GiB 都不一样:

部署 / 配置KV 池token/GiB
canary 改后(0.944 / 256K,capture 到 8)9.01 GiB → 965,793107,191
生产改前(0.91 / 128K,capture 到 32)· 现值7.31 GiB → 679,19192,913
生产改后(0.944 / 256K,capture 到 32,已回滚)7.98 GiB → 855,417107,196
实验 B(0.91 / 256K,已回滚)7.31 GiB → 786,432107,582
实验 A(0.944 / 128K,已回滚)7.98 GiB → 738,76992,578

第一版解释是「它随 capture sizes / max-num-seqs / batched-tokens 变」。但实验 A 把 util 抬到 0.944、 max-model-len 保持 131072,token/GiB 就掉回 92,578(≈基线)。 → 真正的结论:那个 +15.4% 不是 util 带来的、也不是真的多装了 token,它只是 262,144 这个 max_model_len 下的记账口径变化,运行期的激活余量并没有变多 —— Available KV cache memory 反而从 7.98 GiB 掉回 7.31 GiB。 所以「L2 是零显存成本、白解锁 256K」是错的。

事前按「+696 MiB ⇒ +9.3% token」推的是 +63K,实测是 +176K(+25.9%), 当时还把「多给」读成好消息。实际上多给的 token 和崩溃是同一件事的两面: 每 block 装得更多,意味着同样的激活要在更少的余量里跑。 最终口径:不能用任何比例外推,无论另一个部署多「同构」。

回滚

python3 /root/vllm_ctx_tune.py restore vllm-qwen38-27b-lora
# → deployment.apps/vllm-qwen38-27b-lora patched
# → restored ... to args snapshot from /srv/vllm-ctx-tune/vllm-qwen38-27b-lora.prev.json

回滚后复查:args[0] 与 /srv/backup-27b-mem-20260910/vllm-qwen38-27b-lora.yaml 里的 args[0] 逐字节相同(util 0.91 / max-model-len 131072 / max-num-seqs 64 / batched 16384)。 三条回滚路(本次用的是第一条):

python3 vllm_ctx_tune.py restore vllm-qwen38-27b-lora        # 字节级回到 set 之前的 args
kubectl apply -f /srv/backup-27b-mem-20260910/vllm-qwen38-27b-lora.yaml
python3 vllm_ctx_tune.py set vllm-qwen38-27b-lora \
    --gpu-memory-utilization 0.91 --max-model-len 131072     # 反向再 set 一次

一条经验:restore 依赖 /srv/vllm-ctx-tune/<name>.prev.json, 而它只有一个坑位,每次 set 都覆盖。生产这次侥幸没受影响(它只被 set 过一次), canary 的原始配置就丢了。

事后补齐的门禁:longprefill_probe.py

事故的直接教训是缺少一道运行期验收。补上的这道门:

python3 longprefill_probe.py <pod> 60000 80000 124000

已写进手册 §0 铁律 6 与 §5 Phase 1.5:动 util 必须跑它,启动日志三行全绿不算通过。

这次没做,以及为什么

项为什么不
KV 压到 TQ3(推算 +32% token)动精度。serving 侧没有自动质量门禁,必须先有人看输出;且 tq3 部署缺 --kv-cache-dtype-skip-layers,要先对齐才可比。
权重 W4→W3要重跑量化,跌到 4bpw 红线以下,收益还不如 L1+L3。
重开 MTP吞吐杠杆,反向吃显存。2026-08-20「整包提速」里主动去掉的。
删 vision 塔早就没加载(--language-model-only)。
VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0能白拿 0.67 GiB,但等于放弃 capture 阶段的 OOM 保护。
TQ 解码内核 v2已终止不投放。

一个提醒:先看需求,再看容量

改动前拉 9 个生产副本的 /metrics:只有 1 个副本在干活(累计 678 万 prompt token), 其余 8 个 0 running,全集群 KV 池占用 0%。 所以这次升级不是为了当天的吞吐 —— 当前流量下它的收益 ≈ 0,是为将来的长上下文/高并发留余量。 而它把这个「为将来」的余量本身给吃掉了:省出来的 0.67 GiB 正是运行期需要的。 在 KV 池 0% 占用的现网,「给将来多留 176K token」的价值远小于「现在就把 93K 请求打崩」的代价。 → 要动容量/精度之前,先有流量指标或先有质量门禁,两者都没有就别动。