Text to Material
Images are 512 × 512 JPEG previews (quality 90). Four displayed views per object (0, 4, 10, 14), under Aerodynamics Workshop lighting only. Aggregate metrics retain the full original-resolution evaluation across all source scenes, views, and lighting conditions; they are not recomputed on this display subset.
Text-conditioned material comparisons: Ours, VideoMat, and VideoMatGen, rendered with ACES. The same eight displayed objects as the other synthetic viewers, from the full 32-scene evaluation.
Aggregate Rendering Metrics
| Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|
Aggregate Distribution Metrics
| Method | CLIP-FID ↓ | CMMD ↓ | LPIPS-Alex ↓ |
|---|