CLM-019
Claim
1.45B KL-LDM + cross-attn 文本条件在 MS-COCO 达与 Make-A-Scene/GLIDE 相当量级且参数量显著更少。
| id | CLM-019 |
|---|---|
| type | claim |
| source | SRC-2112.10752 |
| scope | LAION-400M 训练; MS-COCO val 评测; 1.45B KL-LDM; BERT tokenizer + τ=transformer; 250 DDIM steps; 最优配置 CFG s=1.5 |
| epistemic | supported |
| lifecycle | RECONCILED |
| created | 2026-09-03 |
| updated | 2026-09-03 |
Proposition
以 cross-attention 文本条件(BERT tokenizer + τ 为 transformer)在 LAION-400M 上训一个 1.45B 参数的 KL-LDM,可在 MS-COCO 上达到与近期最强 AR(Make-A-Scene)与 diffusion(GLIDE)text-to-image 系统相当的量级、且参数量显著更少:开启 CFG(s=1.5)后 FID 12.61 / IS 26.62,低于 GLIDE 12.24 同档、Make-A-Scene 11.84 同档;并显著优于早期 AR(CogView 27.10)与 GAN(LAFITE 26.94)。
Evidence refs
Notes
主表(ms_tables.tex tab:txt2img, 250 DDIM):CogView 27.10/18.20;LAFITE 26.94/26.02;GLIDE 12.24;Make-A-Scene 11.84;LDM-KL-8 23.35/19.93±0.35;LDM-KL-8-G 12.61/26.62±0.38。§4.3.1 原文:"generalizes well to complex, user-defined text prompts"、"on par with the recent state-of-the-art AR and diffusion models ... while substantially reducing parameter count"。注意 MS-COCO val 5K 提示的标准评测协议;参数量对比以"substantially reducing"定性 + 1.45B 具体值为准,不逐模型比参数量(论文未给全对比参数量)。
引用(2)
- EVI-019 主文 T2I 对比表 + §4.3.1:1.45B 文本条件 LDM 对比 AR/diffusion(支撑 CLM-019)。
- SRC-2112.10752 Latent Diffusion Model 原始论文(Rombach et al.);Stable Diffusion 学术源头,DiT 所引『latent diffusion』。
被引用(3)
- EVI-019 主文 T2I 对比表 + §4.3.1:1.45B 文本条件 LDM 对比 AR/diffusion(支撑 CLM-019)。
- FRM-2112.10752 LDM 五层重建:感知压缩 AE + 潜空间扩散的两段分解,换取训练/采样成本大降;多任务多分辨率(256²/512²→~1024²)。
- VER-019 T2I 核验:1.45B LDM-KL-8-G 在 MS-COCO 250 步 FID 12.61/IS 26.62,与 GLIDE 12.24/Make-A-Scene 11.84 同量级(非最低),措辞如实。supported。