Research KB

CLM-008 Claim

3 通道子集 guidance(DiT 实际做法)经 scale 换算后与全 4 通道效果近似:3ch (1+x) ≈ 4ch (1+¾x)。

id CLM-008
type claim
source SRC-2212.09748
scope DiT-XL/2, ImageNet 256x256, 7M steps; FID-50K; guidance on first 3 of 4 latent channels (3ch) vs all 4 channels (4ch)
epistemic supported
lifecycle RECONCILED
created 2026-09-03
updated 2026-09-03

Proposition

对 4 通道 latent 仅施加前 3 通道的 guidance(DiT 论文全部 guidance 实验的实际做法)与全 4 通道 guidance 在调节 scale 后效果近似: 3ch scale (1+x) ≈ 4ch scale (1+¾x);实测 3ch@1.5 → FID-50K 2.27, 4ch@1.375 → 2.20。

Evidence refs

Notes

论文将此现象留作 future work("we leave it to future work to explore this phenomenon further")。注意 4ch@1.375 的 2.20 低于 主表的 2.27,但 Table 2 的正式 SOTA 数字为 3ch 口径——复现 CLM-001 必须按 3 通道施加 guidance。

引用(3)

  • CLM-001 DiT-XL/2(cfg=1.50,3 通道 guidance)在 ImageNet 256×256 达 FID-50K 2.27,为包括 StyleGAN-XL 在内的最低 FID。
  • EVI-008 App. A.1:3 通道 guidance(1.5)FID 2.27 vs 4 通道(1.375)2.20,近似式 3ch(1+x)≈4ch(1+¾x)(支撑 CLM-008)。
  • SRC-2212.09748 DiT 原始论文(Peebles & Xie, ICCV 2023):ViT 替换 LDM U-Net backbone,Gflops 视角的 scaling 研究。

被引用(5)

  • CLM-012 CFG 大幅改善 DiT-XL/2:256 上 FID 9.62→3.22→2.27(无/1.25/1.50),512 上 12.03→4.64→3.04。
  • CPT-003 classifier-free guidance
  • EVI-008 App. A.1:3 通道 guidance(1.5)FID 2.27 vs 4 通道(1.375)2.20,近似式 3ch(1+x)≈4ch(1+¾x)(支撑 CLM-008)。
  • FRM-2212.09748 DiT 五层重建:以 ViT 替换 latent diffusion 的 U-Net backbone、Gflops 为透镜的 scaling 研究;边界为 class-cond ImageNet 256/512。
  • VER-008 核验 CLM-008 subset-channel guidance:App.A.1(3ch@1.5→2.27、4ch@1.375→2.20、近似式)逐字核对一致,'近似'不确定性保留。supported。