perf: cut per-frame CPU cost in the render path - #1
Merged
Merged
Conversation
…leaks Render path (GLView.m drawScene): - Hoist [ctrl getDetail]/getMin/getMax out of the per-pixel depth clamp and replace fmod(i, W) with the loop's column index. The loop ran up to 7 objc_msgSend and 2 fmod calls on each of 307200 pixels, every frame: 4.69 ms -> 0.71 ms per frame. - Swap the mesh index builder to y outer / x inner so it walks the row-major depth buffer sequentially instead of striding W*2 bytes: 0.087 ms -> 0.019 ms at step 1. - Recycle the depth and video frame buffers instead of free()ing and mallocing 600KB/900KB blocks every frame, which crossed the allocator's mmap threshold. Adds -recycleDepthData:/-recycleVideoData:; free() on those buffers stays valid. Both loop rewrites were verified against the originals over 960 randomised parameter combinations: byte-identical buffers, identical mesh index sets and counts. Correctness: - saveSTLB counted facets in an NSUInteger* pointer, so 'faceCount + 2' advanced by 16 and the binary STL header announced 8x the real facet count. Now uint32_t. - binaryVector: allocated NSMutableData without calling init. - All three export paths leaked the 600KB depth buffer; savePly and saveSTL also leaked their accumulator strings. - PLY/STL text built one NSString per coordinate; now a single format pass with byte-identical output. Verified with clang -fsyntax-only against the current macOS SDK. Not built or run against Kinect hardware - the project still targets the 10.6 SDK.
7 tasks
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
Removes the dominant CPU cost from the real-time render loop, and fixes the bugs and leaks found in the export paths along the way.
Why
Three hot spots, all per frame:
The depth clamp called
[ctrl getDetail],[ctrl getMin]and[ctrl getMax]inside the per-pixel loop — up to 7objc_msgSendper pixel — plus twofmod()calls, over 307200 pixels, every frame. The accessors are hoisted and the modulo replaced by the loop's own column index. The export paths inAppDelegate.malready used the ivars directly, so this brings the render path in line with what the author wrote elsewhere.The mesh index builder walked
xouter /yinner, stridingFREENECT_FRAME_W * 2bytes per step through a row-major buffer. Loop order swapped so the walk is sequential.Frame buffers were
freed and re-malloced every frame at 600 KB and 900 KB — above the allocator'smmapthreshold, so each frame paid twommap/munmapround trips plus page faults.-recycleDepthData:/-recycleVideoData:hand them back for reuse; callingfree()on those buffers stays correct, so no existing caller was invalidated.Also fixed here:
saveSTLBcounted facets in anNSUInteger *— a pointer — sofaceCount + 2was pointer arithmetic advancing by 16. The binary STL header announced 8x more facets than were written.binaryVector:allocatedNSMutableDatawithout callinginit.savePlyandsaveSTLalso leaked their accumulator strings.NSStringper coordinate — 12 temporaries per facet. Now a single format pass, byte-identical output.How it was verified
sh scripts/syntax-check.shpassesobjc_msgSendtraffic on an M-series Mac, 200 frames per configuration. First run reported the mesh loop at 0.000 ms because the compiler had eliminated it as dead — the numbers above are from the corrected run with the result consumed.The last two cannot be done here: Xcode 3 project format, 10.6 SDK, 2010 libusb binary, and no Kinect. Reviewers with the hardware should treat the render output as unverified.
Provenance
CONTRIBorAUTHORSfile altered