Skip to content

perf: cut per-frame CPU cost in the render path - #1

Merged
gridboy merged 1 commit into
mainfrom
perf/render-hot-paths
Aug 11, 2026
Merged

gridboy merged 1 commit into
mainfrom
perf/render-hot-paths

Conversation

@gridboy

@gridboy gridboy commented Aug 11, 2026

Copy link
Copy Markdown
Owner

What this changes

Removes the dominant CPU cost from the real-time render loop, and fixes the bugs and leaks found in the export paths along the way.

Why

Three hot spots, all per frame:

before after
depth clamp 4.69 ms 0.71 ms 6.6x
mesh index build (detail 1) 0.087 ms 0.019 ms 4.6x
malloc/free traffic 1.5 MB none recycled

The depth clamp called [ctrl getDetail], [ctrl getMin] and [ctrl getMax] inside the per-pixel loop — up to 7 objc_msgSend per pixel — plus two fmod() calls, over 307200 pixels, every frame. The accessors are hoisted and the modulo replaced by the loop's own column index. The export paths in AppDelegate.m already used the ivars directly, so this brings the render path in line with what the author wrote elsewhere.

The mesh index builder walked x outer / y inner, striding FREENECT_FRAME_W * 2 bytes per step through a row-major buffer. Loop order swapped so the walk is sequential.

Frame buffers were freed and re-malloced every frame at 600 KB and 900 KB — above the allocator's mmap threshold, so each frame paid two mmap/munmap round trips plus page faults. -recycleDepthData: / -recycleVideoData: hand them back for reuse; calling free() on those buffers stays correct, so no existing caller was invalidated.

Also fixed here:

  • saveSTLB counted facets in an NSUInteger * — a pointer — so faceCount + 2 was pointer arithmetic advancing by 16. The binary STL header announced 8x more facets than were written.
  • binaryVector: allocated NSMutableData without calling init.
  • All three export paths leaked their 600 KB depth buffer; savePly and saveSTL also leaked their accumulator strings.
  • PLY/STL text built one NSString per coordinate — 12 temporaries per facet. Now a single format pass, byte-identical output.

How it was verified

  • sh scripts/syntax-check.sh passes
  • Behaviour-preserving changes backed by a comparison against the original: both rewritten loops were run against the originals over 960 randomised parameter combinations (random depth buffers, detail 0-40, min/max sweeps). Byte-identical output buffers; identical index sets and counts for the mesh.
  • Performance claims come with numbers: measured with a benchmark carrying real objc_msgSend traffic on an M-series Mac, 200 frames per configuration. First run reported the mesh loop at 0.000 ms because the compiler had eliminated it as dead — the numbers above are from the corrected run with the result consumed.
  • Built in Xcode
  • Run against real Kinect hardware

The last two cannot be done here: Xcode 3 project format, 10.6 SDK, 2010 libusb binary, and no Kinect. Reviewers with the hardware should treat the render output as unverified.

Provenance

  • No third-party licence header, copyright notice, CONTRIB or AUTHORS file altered
  • No new vendored code

…leaks

Render path (GLView.m drawScene):

- Hoist [ctrl getDetail]/getMin/getMax out of the per-pixel depth clamp and replace
  fmod(i, W) with the loop's column index. The loop ran up to 7 objc_msgSend and 2
  fmod calls on each of 307200 pixels, every frame: 4.69 ms -> 0.71 ms per frame.
- Swap the mesh index builder to y outer / x inner so it walks the row-major depth
  buffer sequentially instead of striding W*2 bytes: 0.087 ms -> 0.019 ms at step 1.
- Recycle the depth and video frame buffers instead of free()ing and mallocing
  600KB/900KB blocks every frame, which crossed the allocator's mmap threshold.
  Adds -recycleDepthData:/-recycleVideoData:; free() on those buffers stays valid.

Both loop rewrites were verified against the originals over 960 randomised parameter
combinations: byte-identical buffers, identical mesh index sets and counts.

Correctness:

- saveSTLB counted facets in an NSUInteger* pointer, so 'faceCount + 2' advanced by
  16 and the binary STL header announced 8x the real facet count. Now uint32_t.
- binaryVector: allocated NSMutableData without calling init.
- All three export paths leaked the 600KB depth buffer; savePly and saveSTL also
  leaked their accumulator strings.
- PLY/STL text built one NSString per coordinate; now a single format pass with
  byte-identical output.

Verified with clang -fsyntax-only against the current macOS SDK. Not built or run
against Kinect hardware - the project still targets the 10.6 SDK.
@gridboy
gridboy merged commit 91d8c80 into main Aug 11, 2026
1 check passed
@gridboy
gridboy deleted the perf/render-hot-paths branch August 11, 2026 13:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant