Low-end hardware (iGPU laptops, Steam Deck): profile and scale fixed frame costs #98

Open
opened 2026-09-27 09:01:29 +00:00 by backspace119 · 2 comments
Owner

Two new testers on low-power hardware see frame rates that barely move with settings:

  • Framework laptop: about 13 fps, fairly constant. Most settings don't change it; only shadows lower it further.
  • Steam Deck (RDNA2 APU, 8 CUs, 1280×800): about 14 fps.

Expected: noticeably more on low settings, and settings that scale the cost down. Flat fps across settings means most of the frame is fixed cost that no setting controls. It's either CPU work, or GPU passes that ignore the preferences.

Step 1: measure (before optimising)

Ask each tester for a log with RenderVKSectionLog + RenderVKGpuTimers on, standing still in a typical scene, at their lowest preset. Plus exact CPU/GPU model (Framework: which mainboard, iGPU or dGPU module), resolution, and whether the laptop is on power.

  • Main-thread frame ≫ GPU total means CPU-bound: scale the per-frame CPU work.
  • GPU total ≈ frame means GPU-bound: scale or skip the fixed passes.

Likely fixed costs (from our own logs on a desktop, which will be several times larger on a laptop CPU)

CPU (ms on a 5090 / 48-core desktop in a busy scene):

  • ui_draw about 6–7: UI rebuilt every frame
  • vk_ingestTick / sweep_total about 4–5
  • vk_syncFromStores about 3
  • texture residency/upload about 2–5
  • LL_objectListUpdate about 2.5
  • avatarCharacter about 1.5
  • vk_rtBuild about 3.5; must be zero when RT is off or unsupported, verify

GPU passes to audit for "runs regardless of settings":

  • SSR
  • reflection probe capture and filtering
  • glow/bloom, SMAA/FXAA, tonemap
  • clustered light binning
  • sky and clouds
  • water
  • depth pre-pass

Candidate work

  1. Settings audit: map every preference to the VK pass it should scale or disable, and wire the ones that are ignored. Presets (Low/Mid) must actually turn things off under VK.
  2. Render scale / upscaling: a render-resolution slider (for example 50–100%) with FSR 1 (spatial, cheap) upscale. This is the biggest single lever on iGPUs and the Deck, and it's standard practice.
  3. CPU budgets that scale with core count and speed: ingest sweep slice, moved queue, texture residency scan, UI redraw only when dirty.
  4. Low-tier pass variants: half-res SSR/probes/bloom, fewer clusters, cheaper lit shader permutation (no RT or projector code paths compiled in).
  5. Frame pacing on the Deck: cap to 30/40 with even pacing, since it can't hold 60 in busy scenes.
  6. Hardware tiering: detect iGPU/APU and pick defaults per tier. Keep "best features the hardware supports" as the rule, not the lowest common denominator.

Tracking ticket; split into sub-issues once the logs say where the time goes.

Two new testers on low-power hardware see frame rates that barely move with settings: - **Framework laptop:** about 13 fps, fairly constant. Most settings don't change it; only shadows lower it further. - **Steam Deck** (RDNA2 APU, 8 CUs, 1280×800): about 14 fps. Expected: noticeably more on low settings, and settings that scale the cost down. Flat fps across settings means most of the frame is **fixed cost** that no setting controls. It's either CPU work, or GPU passes that ignore the preferences. ### Step 1: measure (before optimising) Ask each tester for a log with `RenderVKSectionLog` + `RenderVKGpuTimers` on, standing still in a typical scene, at their lowest preset. Plus exact CPU/GPU model (Framework: which mainboard, iGPU or dGPU module), resolution, and whether the laptop is on power. - Main-thread `frame` ≫ GPU total means CPU-bound: scale the per-frame CPU work. - GPU total ≈ frame means GPU-bound: scale or skip the fixed passes. ### Likely fixed costs (from our own logs on a desktop, which will be several times larger on a laptop CPU) CPU (ms on a 5090 / 48-core desktop in a busy scene): - `ui_draw` about 6–7: UI rebuilt every frame - `vk_ingestTick` / `sweep_total` about 4–5 - `vk_syncFromStores` about 3 - texture residency/upload about 2–5 - `LL_objectListUpdate` about 2.5 - `avatarCharacter` about 1.5 - `vk_rtBuild` about 3.5; must be zero when RT is off or unsupported, verify GPU passes to audit for "runs regardless of settings": - SSR - reflection probe capture and filtering - glow/bloom, SMAA/FXAA, tonemap - clustered light binning - sky and clouds - water - depth pre-pass ### Candidate work 1. **Settings audit:** map every preference to the VK pass it should scale or disable, and wire the ones that are ignored. Presets (Low/Mid) must actually turn things off under VK. 2. **Render scale / upscaling:** a render-resolution slider (for example 50–100%) with FSR 1 (spatial, cheap) upscale. This is the biggest single lever on iGPUs and the Deck, and it's standard practice. 3. **CPU budgets that scale with core count and speed:** ingest sweep slice, moved queue, texture residency scan, UI redraw only when dirty. 4. **Low-tier pass variants:** half-res SSR/probes/bloom, fewer clusters, cheaper lit shader permutation (no RT or projector code paths compiled in). 5. **Frame pacing on the Deck:** cap to 30/40 with even pacing, since it can't hold 60 in busy scenes. 6. **Hardware tiering:** detect iGPU/APU and pick defaults per tier. Keep "best features the hardware supports" as the rule, not the lowest common denominator. Tracking ticket; split into sub-issues once the logs say where the time goes.
Author
Owner

Data point: Steam Deck goes from about 14 to about 40 fps with shadows off. That's about 71 ms → 25 ms per frame, so shadows cost about 45 ms on the Deck, roughly two thirds of the frame. The rest of the frame (25 ms) is still above 60 fps, but within reach of 30/40 fps with some trimming.

What the code does today (VK sun shadows):

  • 4 cascades × 2048² D32 (NUM_CASCADES = 4, RenderVKShadowResolution = 2048). This is fixed: no preference or graphics preset changes the resolution or the cascade count. LL's RenderShadowResolutionScale isn't read by the VK path, and none of this is in the feature table.
  • So the only shadow control a user has is all-or-nothing (Shadows: None / Sun / Sun + Projectors). That matches the testers' "only shadows change anything".

Shadow-specific items for this ticket (standard practice in modern engines):

  1. Wire resolution to prefs/presets: Low/Mid, for example 1024², High 2048², Ultra 4096², and honour RenderShadowResolutionScale or replace it.
  2. Cascade count by tier: 2 cascades on low tiers (near + far), 3 on mid, 4 on high.
  3. Cache and stagger the far cascades: re-render cascade 0 every frame, 1 every 2nd, and 2 and 3 every 4th/8th (or only when the sun or casters in them change), with texel-snapped stable projections so the cache stays valid. Most static-world shadow cost disappears.
  4. Caster culling / LOD for shadows: cheaper LOD and small-caster rejection for the far cascades (RenderVKShadowMinCasterTexels exists; check its default for low tiers).
  5. Get the shadow pass split out in the GPU timer log (RenderVKGpuTimers) per cascade so these can be measured on the Deck.

Also: check which shadow setting the Deck tester had (Sun only, or Sun + Projectors). If projectors were on, part of the 45 ms may be projector tiles; the new projector budget slider (not yet released) covers that side.

**Data point: Steam Deck goes from about 14 to about 40 fps with shadows off.** That's about 71 ms → 25 ms per frame, so **shadows cost about 45 ms on the Deck**, roughly two thirds of the frame. The rest of the frame (25 ms) is still above 60 fps, but within reach of 30/40 fps with some trimming. What the code does today (VK sun shadows): - **4 cascades × 2048² D32** (`NUM_CASCADES = 4`, `RenderVKShadowResolution` = 2048). This is fixed: no preference or graphics preset changes the resolution or the cascade count. LL's `RenderShadowResolutionScale` isn't read by the VK path, and none of this is in the feature table. - So the only shadow control a user has is all-or-nothing (Shadows: None / Sun / Sun + Projectors). That matches the testers' "only shadows change anything". Shadow-specific items for this ticket (standard practice in modern engines): 1. **Wire resolution to prefs/presets:** Low/Mid, for example 1024², High 2048², Ultra 4096², and honour `RenderShadowResolutionScale` or replace it. 2. **Cascade count by tier:** 2 cascades on low tiers (near + far), 3 on mid, 4 on high. 3. **Cache and stagger the far cascades:** re-render cascade 0 every frame, 1 every 2nd, and 2 and 3 every 4th/8th (or only when the sun or casters in them change), with texel-snapped stable projections so the cache stays valid. Most static-world shadow cost disappears. 4. **Caster culling / LOD for shadows:** cheaper LOD and small-caster rejection for the far cascades (`RenderVKShadowMinCasterTexels` exists; check its default for low tiers). 5. Get the shadow pass split out in the GPU timer log (`RenderVKGpuTimers`) per cascade so these can be measured on the Deck. Also: check which shadow setting the Deck tester had (Sun only, or Sun + Projectors). If projectors were on, part of the 45 ms may be projector tiles; the new projector budget slider (not yet released) covers that side.
Author
Owner

Datapoint: GTX 1080 Ti + i7-6700K, Windows, 1918×1008, build 946e89a915 (RenderVKGpuTimers log). GPU-bound throughout.

State GPU ms/frame FPS
Shadows off ~40 (pass1 21, pass2 8.7, probe 3.4–4.2, sel+hud 3.1–3.4, compute 2) ~23–25
Cascades on (4×2048), no projectors ~71–81 (c_shadows 33–38) ~13–15
Cascades + projector shadows, proj_n=13–16 (default budget 16) ~118–136 (c_projshadow 42–65) ~6–8

Pass1 breakdown (shadows off):

Sub-pass ms
g_prepass (depth prepass) 6.8
g_pass1 14.2
g_to_glow 9.5
g_passes 3.1

For scale: the same scene class on an RTX 5090 was about 7 ms total. The 1080 Ti is roughly 6–9× slower, so our frame cost is budgeted for high-end cards.

Actions:

  1. Presets must scale shadows:
    • The projector budget defaults to 16 on every preset; it should be something like 0 / 2 / 8 / 16 by preset.
    • Cascade resolution and "Shadow updates" (cached / distant-only) should also be set per preset.
    • Auto-pick the preset tier from the GPU: VRAM, device class, or a first-run GPU timing sample.
  2. Geometry cost:
    • The depth prepass plus pass1 add up to about 21 ms, which suggests the frame is vertex-bound.
    • Add a mesh LOD bias for low tiers.
    • Shadow LOD: a lower LOD for shadow casters, and cull small casters by their projected size in each cascade.
    • Consider skipping the prepass on low tiers.
  3. Probe capture (~3.5 ms/frame): time-slice it to one face per frame. This is already in the backlog.
  4. Investigate sel+hud: it costs a constant ~3 ms with nothing selected.
  5. Investigate g_to_glow: break down its ~9.5 ms (alpha/transparent work).
**Datapoint: GTX 1080 Ti + i7-6700K, Windows, 1918×1008, build 946e89a915 (`RenderVKGpuTimers` log).** GPU-bound throughout. | State | GPU ms/frame | FPS | |---|---|---| | Shadows off | ~40 (pass1 21, pass2 8.7, probe 3.4–4.2, sel+hud 3.1–3.4, compute 2) | ~23–25 | | Cascades on (4×2048), no projectors | ~71–81 (`c_shadows` 33–38) | ~13–15 | | Cascades + projector shadows, `proj_n`=13–16 (default budget 16) | ~118–136 (`c_projshadow` 42–65) | ~6–8 | **Pass1 breakdown (shadows off):** | Sub-pass | ms | |---|---| | `g_prepass` (depth prepass) | 6.8 | | `g_pass1` | 14.2 | | `g_to_glow` | 9.5 | | `g_passes` | 3.1 | For scale: the same scene class on an RTX 5090 was about 7 ms total. The 1080 Ti is roughly 6–9× slower, so our frame cost is budgeted for high-end cards. **Actions:** 1. **Presets must scale shadows:** - The projector budget defaults to 16 on every preset; it should be something like 0 / 2 / 8 / 16 by preset. - Cascade resolution and "Shadow updates" (cached / distant-only) should also be set per preset. - Auto-pick the preset tier from the GPU: VRAM, device class, or a first-run GPU timing sample. 2. **Geometry cost:** - The depth prepass plus pass1 add up to about 21 ms, which suggests the frame is vertex-bound. - Add a mesh LOD bias for low tiers. - Shadow LOD: a lower LOD for shadow casters, and cull small casters by their projected size in each cascade. - Consider skipping the prepass on low tiers. 3. **Probe capture (~3.5 ms/frame):** time-slice it to one face per frame. This is already in the backlog. 4. **Investigate `sel+hud`:** it costs a constant ~3 ms with nothing selected. 5. **Investigate `g_to_glow`:** break down its ~9.5 ms (alpha/transparent work).
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
backspace119/Slipstream#98
No description provided.