Research KB

CLM-009 Claim

U-Net 归纳偏置对 diffusion 质量非关键:标准 ViT 可直接替换 U-Net backbone 并取得更优质量。

id CLM-009
type claim
source SRC-2212.09748
scope class-conditional ImageNet 256x256 & 512x512 latent diffusion; DiT (standard ViT backbone) vs U-Net baselines (ADM/LDM) under ADM evaluation suite
epistemic partially-supported
lifecycle RECONCILED
created 2026-09-03
updated 2026-09-03

Proposition

U-Net 的归纳偏置对 diffusion 模型的性能不是关键:标准 ViT 架构 (沿用 ViT 惯例、不含卷积 U-Net 组件)可在 latent diffusion 中 直接替换 U-Net backbone,并在同等评估协议下取得更优的生成质量。

Evidence refs

Notes

论文的设计论题(§1:"the U-Net inductive bias is not crucial to the performance of diffusion models")。经验支撑即 CLM-001/004 的 SOTA 对比(DiT 优于全部既有 U-Net diffusion 模型);本 claim 将 对比结果提炼为架构层论断。限定:证据限于 class-conditional latent diffusion 两个分辨率;pixel-space 与 text-to-image 设定未验证 (结论段以 "could be explored" 表述为 future work)。

引用(3)

  • CLM-001 DiT-XL/2(cfg=1.50,3 通道 guidance)在 ImageNet 256×256 达 FID-50K 2.27,为包括 StyleGAN-XL 在内的最低 FID。
  • EVI-009 §1 + Table 2/3/6:DiT-XL/2 2.27(256)/3.04(512)vs LDM-4-G 3.60/ADM-G+U 3.94 及 118.6G/524.6G 算力对照——架构论题的经验支撑(支撑 CLM-009)。
  • SRC-2212.09748 DiT 原始论文(Peebles & Xie, ICCV 2023):ViT 替换 LDM U-Net backbone,Gflops 视角的 scaling 研究。

被引用(4)

  • CFL-001 表面冲突:CLM-009(DiT)称'U-Net 归纳偏置对生成质量非关键,可换 ViT',CLM-015(LDM)称'2D 卷积归纳偏置对重建保真关键,容忍低压缩'。论文核实:两处'U-Net'是不同网络(second-stage 去噪 UNet vs first-stage 压缩 autoencoder),同名异物、可同真互补,提议 terminology-difference。
  • EVI-009 §1 + Table 2/3/6:DiT-XL/2 2.27(256)/3.04(512)vs LDM-4-G 3.60/ADM-G+U 3.94 及 118.6G/524.6G 算力对照——架构论题的经验支撑(支撑 CLM-009)。
  • FRM-2212.09748 DiT 五层重建:以 ViT 替换 latent diffusion 的 U-Net backbone、Gflops 为透镜的 scaling 研究;边界为 class-cond ImageNet 256/512。
  • VER-009 核验 CLM-009 架构论题:可检验部分(ViT 替换 U-Net 质量反升)与 Table 2/3/6 一致;'归纳偏置非关键'属作者诠释(非受控对比),按 A4 保留。partially-supported。