Text to Material

Images are 512 × 512 JPEG previews (quality 90). Four displayed views per object (0, 4, 10, 14), under Aerodynamics Workshop lighting only. Aggregate metrics retain the full original-resolution evaluation across all source scenes, views, and lighting conditions; they are not recomputed on this display subset.

Text-conditioned material comparisons: Ours, VideoMat, and VideoMatGen, rendered with ACES. The same eight displayed objects as the other synthetic viewers, from the full 32-scene evaluation.

Visible methods
Visible lighting
Aggregate Rendering Metrics
MethodPSNR ↑SSIM ↑LPIPS ↓
Aggregate Distribution Metrics
MethodCLIP-FID ↓CMMD ↓LPIPS-Alex ↓