NVIDIA NVIDIA Real-Time Graphics Research NVIDIA Research

Texture Space Material Diffusion

High quality PBR materials from photos, text, or low-resolution maps — generated entirely in texture space

NVIDIA
Figure 1: Teaser

Materials from photos, material maps in 8K resolution, materials from text, and neural materials. Our texture space diffusion model generates materials from image or text conditioning.

Abstract


We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, enabling the diffusion process to generalize across arbitrary geometries and texture parameterizations. This approach also avoids the view consistency issues inherent in video and multi-view diffusion models. Because texture space is two dimensional, we can reuse the strong priors of pretrained video diffusion models. We apply our method to high quality material reconstruction from posed photos captured under unknown lighting, as well as to text- and image guided material generation. Our method can scale to high resolutions (8K), 100+ input views, and neural material representations. In quantitative and qualitative evaluations we show state-of-the-art results for material generation and reconstruction.

Why texture space?


Image-space methods generate multiple views and reconstruct materials from them, but the same surface point may be predicted differently across views, leading to blur, lost detail and suppressed highlights. Object-space methods (sparse voxels, triplanes) are view consistent but memory hungry and trained from scratch. Texture space offers the best of both: it is view consistent by design and two-dimensional, so powerful pretrained image and video diffusion priors can be reused directly. TEXGen (Yu et al., 2024) showed that it is possible to directly generate albedo maps in texture space. Inspired by this, we formulate a texture space generation process for complete PBR materials.

grid_on

Material generation in texture space

To our knowledge, the first diffusion framework that generates complete PBR materials (base color, height, roughness, metalness) entirely in texture space.

photo_camera

Robust under unknown lighting

Demodulates lighting from single or multiple views — photos, phone captures, or frames from a video model — and inpaints unobserved regions.

zoom_out_map

Scales at inference time

Training-free noise rolling and coverage-aware expert aggregation scale to 8K textures and over 100 input views.

Texture space observations


Figure 2: Texture space observations

Geometry, parameterization, multi-view captures, and texture space observations. Given camera poses and geometry with a non-overlapping UV parameterization, each view is projected back into texture space, producing partially covered observations. Formulating diffusion in this 2D domain lets us leverage powerful image and video diffusion priors.

Method


Figure 3: Method overview

Method overview. A finetuned DiT video model operates on texture space inputs. In all modes the model is conditioned on texture space world positions, normals and a text prompt. It demodulates the unknown lighting and produces clean albedo, roughness and metallicity maps with high frequency detail, ready for standard content creation tools.

We finetune Wan 2.1-1.3B with a flow matching objective. The texture space views, the shared conditions (world space positions and normals), and the noisy material latents are concatenated along the temporal dimension, so the transformer can attend across all of them. The output is two texture space images: base color, and an RGB-packed height / roughness / metallicity map. Training uses 120k path-traced objects from TexVerse, 17 frames per object, with resolution progressively increased from 512² to 2048².

Multi-view → material17 shaded texture space views are merged into dense material maps. Training-time FLUX img2img and blur augmentations make the model robust to imperfect views.
Single-view → materialOne sparse observation; the model demodulates lighting and inpaints the rest, guided by a novel 3D-aware RoPE.
Text → materialThe input view is replaced with a UV coverage mask; the 3D-aware RoPE supplies geometric structure.
UpscalingA generative refiner conditioned on low-resolution PBR maps, combined with noise rolling to reach 8K.

Inference scaling: 8K textures and 100+ views

Progressive noise rolling. We generate at 2K, upscale to 8K, add a small amount of noise and resume denoising on non-overlapping 2K crops, randomly rolling the crop grid each step to hide seams. Coverage-aware expert aggregation. Input views are split into overlapping batches; each batch predicts the flow for the same noisy latent. Per texel, we keep the top-2 experts by view coverage and average them weighted by coverage, keeping redundancy without over-smoothing.

Figure 8: Inference scaling ablation

17 views at 2K vs. 117 views at 8K. Our inference augmentations capture high-frequency detail and more accurate specular predictions while preserving fidelity (28.42 dB / 0.919 / 0.0668 → 28.34 dB / 0.917 / 0.0639 PSNR / SSIM / LPIPS). Asset from the Digital Twin Catalog dataset.

Material reconstruction from real photos


Figure 4: Real-world reconstruction

Reconstruction from posed photographs of two objects from the DTC dataset, shown under the capture lighting and two novel lighting conditions. DiffPT (differentiable path tracing) bakes lighting into the materials, while LSRM jointly reconstructs geometry and materials. Our method yields clean, relightable materials. Assets from the Digital Twin Catalog dataset.

MethodReal recon (136 img)Synthetic recon (256 img)Synthetic relight (2k img)
PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
Known (reference) geometry
DiffPT30.670.9590.030727.500.9290.073924.900.9030.0911
Ours29.030.9470.033131.240.9600.037030.460.9500.0401
Reconstructed geometry (LSRM)
LSRM27.750.9330.038822.820.8720.121021.820.8620.1241
Ours (LSRM geo)27.010.9360.042324.100.8900.089623.710.8850.0927

Multi-view to materials from 17 views at 2K. DiffPT overfits the capture lighting (best real reconstruction) but fails to generalize under novel lighting, where our method leads by over 5 dB.

Generative materials


Single image → material

Figure 6: Single-view generation

Materials conditioned on a single image and text description, compared with Hunyuan3D 2.1, VideoMatGen and Trellis.2 using the same reference geometry and conditioning view. The visualized view is rotated slightly from the conditioning view (leftmost insets).

Single-view → materialCLIP-FID↓CMMD↓LPIPS↓Text → materialCLIP-FID↓CMMD↓LPIPS↓
Hunyuan3D 2.13.4190.04370.0520
VideoMatGen2.9730.02360.0527VideoMatGen4.7250.03300.0639
Trellis.22.2270.01840.0395VideoMat4.3760.02470.0652
Ours1.5200.00810.0325Ours3.3380.02300.0542

All metrics evaluated on 32 scenes from from the BlenderVault test set, using 8 views × 8 light probes for each scene.

Text → material

Figure S9: Text to material

Materials generated from a text prompt and the input geometry only (UV mask, world space positions and normals), compared with VideoMat and VideoMatGen. Without any image guidance, the 3D-aware RoPE gives enough structure to produce semantically meaningful materials.

Neural materials. Replacing the PBR target latents with VAE-encoded neural material textures extends the image-conditioned model to richer appearance, such as the fuzzy material in the right column of the teaser.

3D-aware rotary positional embedding

Neighboring surface points can be far apart in texture space. Inspired by RomanTex , we replace Wan's RoPE over pixel coordinates and frame id (puv+fid) with a 3D-aware RoPE over the texel's world space position and frame id (pxyz+fid). Attention then concentrates on regions that are close in 3D, even when they are distant in the UV atlas.

Figure S6: Attention visualization

Attention activations for a query point (red circle) at five layers. Top: Wan 2.1 RoPE (puv+fid). Middle: Wan 2.1 RoPE with G-buffer position and normal guides. Bottom: our 3D-aware RoPE with G-buffer guides, which attends strongly to points that are close in world space.

RoPE variantInputsG-bufferCLIP-FID↓CMMD↓LPIPS↓
Wan 2.1puv+fid✗4.0360.04050.1005
Wan 2.1puv+fid✓3.6770.02890.0971
3D-aware (ours)pxyz+fid✓3.5610.02370.0957

RoPE ablation on text → material generation (32 scenes × 8 views × 8 probes, 512² textures). The 3D-aware RoPE gives a small but consistent improvement on all metrics.

Material components


Rendered-image metrics measure the final appearance, but not whether each material parameter is recovered correctly. We therefore evaluate the individual components — base color, roughness and metallicity — of materials reconstructed from 17 views. Because LSRM reconstructs its own geometry, and hence a different UV space, the comparison is made on g-buffer renderings of the material maps rather than on the texture images themselves: 17 views of each of the 32 scenes in the BlenderVault test set.

Figure S2: Material components

Figure S2: Material components. G-buffer renderings of base color (adjusted for intensity), roughness and metallicity for DiffPT, LSRM, our method and the reference. Our maps stay closest to the reference on all three components, with a clean base color and roughness that follows the surface detail.

Reconstructed base color carries a systematic intensity shift, since observed brightness depends jointly on illumination and reflectance. To compensate, we additionally report scale-invariant PSNR.

MethodBase colorRoughnessMetallicity
siPSNR↑PSNR↑PSNR↑PSNR↑
Ours29.1927.0024.3218.24
DiffPT25.0820.2117.9813.23
LSRM23.1715.4613.1814.73

Table S3: Material components. Metrics on g-buffer renderings, 32 scenes × 17 views, from multi-view guidance. siPSNR is the scale-invariant PSNR of the base color; the remaining columns are standard PSNR.

Limitations


Inference is not yet optimized: the 17-view model takes about 130 s at 2K on a GB300 GPU. Capturing specular appearance remains challenging. Wan 2.1 attends over 16×16 patches, so patches spanning several UV atlas segments cannot be fully disambiguated; patch-aware unwrapping or pixel-space diffusion are promising directions.

Citation


@article{munkberg2026texdiff,
    author  = {Jacob Munkberg and Peter Kocsis and Jon Hasselgren},
    title   = {{Texture Space Material Diffusion}},
    journal = {arXiv preprint arXiv:2609.37654},
    year    = {2026}
}

Paper


Paper preview