Models / sdxl-base-unet

SDXL Base UNet

The 2.6B-parameter denoising UNet of Stable Diffusion XL base: residual blocks and spatial transformers at three resolutions, cross-attending to a 2048-wide text embedding, conditioned on the timestep and on SDXL's pooled-text-plus-micro-conditioning vector.

2.6B parameterssdxlCreativeML Open RAIL++-Mimage-generationvisionconvolutionaldiffusiontext-to-imageunet
Prompt

A red fox curled up asleep on a mossy boulder in a misty pine forest at dawn, soft golden light through the trees, dew on the moss, detailed fur, shallow depth of field, wildlife photograph

Generated with the Linnet UNet
Linnet, Linnet torch (generated source), bf16, cuda
Generated with the stock pipeline
diffusers StableDiffusionXLPipeline (stock UNet), bf16

StableDiffusionXLPipeline, 25 steps, 1024x1024, seed 1234, guidance 5.0: once with this card's UNet (Linnet torch) in place of the pipeline's own, once stock. The text encoders, scheduler and VAE are the pipeline's in both.

PSNR, Linnet vs diffusers
27.12 dB
Wall time, Linnet
4.74 s
Wall time, diffusers
1.27 s