Skip to content

WebGLRenderer: with a transmissive object in view, every opaque material looks its program up twice a frame #34654

Description

@sertacfirat

Description

When a scene with a transmissive material is rendered to the canvas, renderTransmissionPass() draws the opaque objects into the transmission render target, and the main pass then draws them onto the canvas. The two passes use different output settings: the render target renders in the working colour space with NoToneMapping, and the canvas uses renderer.outputColorSpace (sRGB by default) and renderer.toneMapping. So at each switch setProgram() finds that materialProperties.outputColorSpace !== colorSpace (and materialProperties.toneMapping !== toneMapping when tone mapping is on), and calls getProgram() for every material drawn in both passes, twice a frame. Each call runs getParameters() and getProgramCacheKey(), finds the program already compiled in materialProperties.programs, and resets uniformsList, so the next upload rebuilds the list with WebGLUniforms.seqWithValue().

Nothing gets compiled, but the lookups cost main-thread CPU time on every frame, in proportion to the number of materials. This happens with the renderer's default settings too, because the sRGB output alone differs from the render target. With post-processing it doesn't happen, because the main pass also renders into a render target, so both passes agree.

We found this in a product page built on a CAD model. On the tablet profile, which has no composer, the lookups took 36% of the page's JS time in a CPU profile of a camera drag and a scroll, and the main thread of a laptop with integrated Intel Arc graphics was saturated.

Reproduction steps

  1. Open the page below and read the title: render()'s CPU time and the number of program lookups per frame. The lookups are counted through customProgramCacheKey(), which getParameters() calls once per lookup.
  2. Comment out the transmissive sphere: the lookups drop to 0.
  3. Or keep the sphere and set renderer.outputColorSpace = THREE.LinearSRGBColorSpace with NoToneMapping: the passes then agree, and the lookups drop to 0 as well.

Code

<script type="importmap">
{ "imports": { "three": "https://cdn.jsdelivr.net/npm/three@0.186.1/build/three.module.js" } }
</script>
<script type="module">
import * as THREE from 'three';

const renderer = new THREE.WebGLRenderer( { antialias: true } );
renderer.setSize( window.innerWidth, window.innerHeight );
renderer.toneMapping = THREE.ACESFilmicToneMapping; // the defaults (sRGB output, NoToneMapping) switch programs as well
document.body.appendChild( renderer.domElement );

const scene = new THREE.Scene();
scene.add( new THREE.HemisphereLight( 0xffffff, 0x404040, 2 ) );
const camera = new THREE.PerspectiveCamera( 45, window.innerWidth / window.innerHeight, 0.1, 100 );
camera.position.z = 16;

// getParameters() calls customProgramCacheKey() once per program lookup: count them
let lookups = 0;
const key = THREE.Material.prototype.customProgramCacheKey;

const geometry = new THREE.BoxGeometry( 0.4, 0.4, 0.4 );

for ( let i = 0; i < 300; i ++ ) {

	const material = new THREE.MeshStandardMaterial( { color: new THREE.Color().setHSL( i / 300, 0.6, 0.5 ) } );
	material.customProgramCacheKey = function () { lookups ++; return key.call( this ); };
	const mesh = new THREE.Mesh( geometry, material );
	mesh.position.set( ( i % 20 ) - 9.5, Math.floor( i / 20 ) - 7, - 3 );
	scene.add( mesh );

}

// one transmissive object
scene.add( new THREE.Mesh( new THREE.SphereGeometry( 2, 64, 32 ), new THREE.MeshPhysicalMaterial( { transmission: 1, roughness: 0.1 } ) ) );

renderer.setAnimationLoop( () => {

	lookups = 0;
	const t0 = performance.now();
	renderer.render( scene, camera );
	document.title = `render() ${ ( performance.now() - t0 ).toFixed( 1 ) } ms, ${ lookups } program lookups`;

} );
</script>

Measurements

We measured a setup like the one above: 300 boxes with their own MeshStandardMaterial and ACES tone mapping, 1280 x 720, Chrome 154 on Windows, Ryzen 9 7950X with an RTX 4070 Ti. Each figure is the median CPU time of render() over 600 frames.

render() CPU throttled 4x lookups / frame
no transmissive object 0.5 ms 2.8 ms 0
transmissive sphere 4.7 ms 28.3 ms 600
transmissive sphere, canvas in linear + NoToneMapping (no switch) 1.2 ms 4.5 ms 0
transmissive sphere, with the change sketched below 1.3 ms 5.5 ms 0

With the renderer's defaults (sRGB output, NoToneMapping), the transmissive sphere case measured 5.0 ms and 600 lookups.

A possible fix

We patched r186 locally, and we can open a PR if this direction is welcome:

  • Take the two output checks out of the else if chain in setProgram(). Compare them only when the rest of the chain found no difference.
  • When only the output colour space or the tone mapping differs, look up a small per-material map from colorSpace + ':' + toneMapping to the program found earlier. Switch to that program instead of calling getProgram(). Keep its uniform list while materialProperties.uniforms is still the object the list was built from.
  • Clear the map whenever getProgram() runs for any other reason: the material's version, the lights state, the object's features, the environment, the fog or the clipping. A remembered program then always matches the rest of the current state.
  • Leave node materials on the existing path.

We first tried keeping the program while the object stayed the one it was found for. That rendered three cases wrong:

  • a material drawn in two scenes with different lights;
  • a normal-mapped material shared by a mesh with tangents and one without;
  • instance colours added after the first frame.

The version above renders these three cases, and two more, identically to unpatched r186, frame by frame. We have a small pixel-comparison page for them and can include it with the PR.

#28423 would remove the second draw of the opaque objects altogether, and with it this switch. This fix is much smaller and covers the case where the opaque pass is still rendered twice. It is also related to #22530, where program lookups repeated every frame for materials shared by skinned and non-skinned meshes.

Version

r186 (dev at b745e6c has the same code)

Device

Desktop

Browser

Chrome

OS

Windows

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions