Skip to content

2/5 — internal/mix: mixBlock(), the whole signal path in one function #3

Description

@KevinGruber2001

Part 2 of #1. Depends on #2. Still no device, still no malgo, still nothing deleted — this builds the mixer alongside the existing engine and proves it by rendering a file and comparing.

Problem

ma_engine mixes for us today. Groups sum their children, ma_node chains apply effects, and Go sets values from outside. The consequences:

  • We cannot insert our own code anywhere in the signal path
  • Live playback and offline render take different paths through the same graph, and already produce different output — automation runs on a 20 ms wall-clock ticker live (automationLoop in internal/engine/automation.go) but samples once per 4096-frame chunk in Render(), which is ~85 ms at 48 kHz
  • Nothing about mixing is testable without an audio device

This issue replaces all of it with one function you can read.

Before you start

  • Read internal/engine/graph.go first. It builds clip → track → bus → master out of miniaudio groups. The new mixer builds the same tree; it just does the summing itself. Knowing the existing shape makes the new code obvious.
  • Mixing is addition. Summing two tracks is out[i] += a[i] + b[i]. There is no cleverness here at all. If a line of this mixer looks clever, it's probably wrong.
  • Gain in dB → linear is 10^(dB/20). -6 dB ≈ 0.501. miniaudio was doing this for you via ma_volume_db_to_linear.
  • Constant-power panning. Linear panning (L = 1-p, R = p) makes the centre sound quieter than the sides, because power is amplitude squared. Constant power uses a quarter-circle: angle = (pan+1) * π/4, L = cos(angle), R = sin(angle). Centre comes out at ≈0.707 (−3 dB) on both sides. Pick this now — changing it later changes how every existing project sounds.
  • Reread clipFilePosition in internal/engine and its test. That logic (where in the source file are we, accounting for the in-point and looping?) transfers over directly — you already solved it.

Proposed solution

flowchart TD
  subgraph "one call to MixBlock()"
    Z["zero the buffers"] --> T
    T["for each track:<br/>read its clips into trackBuf"] --> G["apply gain · pan · mute/solo"]
    G --> S["sum into its bus buffer"]
    S --> B["for each bus:<br/>sum into master"]
    B --> M["master gain"] --> O["out []float32"]
  end
  A["automation<br/>sampled once, here"] -.-> G
Loading

One function, called with a slice to fill. Whoever calls it decides what the audio is for:

flowchart LR
  W["WAV writer<br/>(render)"] -- pulls --> MB["MixBlock()"]
  D["malgo device<br/>(playback)"] -- pulls --> MB
Loading

That is the whole reason render and playback can never drift apart again — there is only one implementation.

Draft code

// Package mix turns a loaded project into audio, one block at a time.
package mix

// Mixer renders a project to audio. Create one with New, then call MixBlock
// repeatedly — each call produces the next chunk of the timeline.
//
// A Mixer is NOT safe for concurrent use: exactly one goroutine calls
// MixBlock. (The device callback doesn't — see issue 4, the ring buffer.)
type Mixer struct {
    proj  *project.Project // immutable snapshot from state.Store.Get()
    rate  int
    chans int              // always 2 for now

    pos   uint64           // playhead, in frames. THE source of truth.
    clips []clipState      // runtime read positions, one per clip

    // Scratch buffers, allocated once in New and reused forever.
    // MixBlock must never allocate.
    trackBuf []float32
    busBufs  map[string][]float32
    masterBuf []float32
}

// MixBlock fills out with the next len(out)/channels frames and advances the
// playhead. out is interleaved stereo float32.
//
// Read this top to bottom and you have read the entire signal path.
func (m *Mixer) MixBlock(out []float32) {
    frames := len(out) / m.chans
    beat := m.beatAt(m.pos) // sample automation ONCE per block, from the frame clock

    zero(out)
    for _, buf := range m.busBufs {
        zero(buf)
    }

    for ti, t := range m.proj.Tracks {
        zero(m.trackBuf)

        // 1. clips -> trackBuf
        for ci := range t.Clips {
            m.readClip(ti, ci, m.trackBuf, frames)
        }

        // 2. (issue 3 inserts the FX chain here)

        // 3. gain + pan + mute/solo -> the track's bus
        gain := m.trackGain(t, beat)     // dB -> linear, automation applied
        l, r := panGains(m.trackPan(t, beat))
        dst := m.busBufs[t.Bus]          // "" == straight to master
        for f := 0; f < frames; f++ {
            dst[f*2]   += m.trackBuf[f*2]   * gain * l
            dst[f*2+1] += m.trackBuf[f*2+1] * gain * r
        }
    }

    // 4. buses -> master
    for _, b := range m.proj.Buses {
        g := db2lin(b.Gain)
        src := m.busBufs[b.ID]
        for i := range m.masterBuf {
            m.masterBuf[i] += src[i] * g
        }
    }

    // 5. master -> out
    g := db2lin(m.proj.Master.Gain)
    for i := range out {
        out[i] = m.masterBuf[i] * g
    }

    m.pos += uint64(frames)
}

Reading a clip is where the only real bookkeeping lives — and note the fade, which is not optional:

// readClip mixes the part of one clip that falls inside this block into dst.
//
// The 5 ms fade at each edge is mandatory, not a nicety: a clip that starts
// or ends mid-waveform produces a step change in the signal, which is a click.
// Since trimming and splitting are features, every clip has cut edges.
const fadeFrames = 240 // 5 ms @ 48 kHz

func (m *Mixer) readClip(trackIdx, clipIdx int, dst []float32, frames int) {
    // - is the clip active anywhere in [pos, pos+frames)?
    // - where in the source file does that correspond to (offset + loop wrap)?
    // - copy, applying clip gain and the edge fades
}

And the panning, written out because it's the one formula worth getting right once:

// panGains converts a pan position (-1 left .. +1 right) into left and right
// multipliers using a constant-power law: centre is -3 dB on both sides rather
// than full level, so sweeping across the stereo field doesn't get louder in
// the middle.
func panGains(pan float64) (l, r float32) {
    angle := (pan + 1) * math.Pi / 4 // -1..1  ->  0..π/2
    return float32(math.Cos(angle)), float32(math.Sin(angle))
}

func db2lin(db float64) float32 { return float32(math.Pow(10, db/20)) }

Proving it

The whole point of doing this before touching the device: you can check your work against the engine you already trust.

go run ./cmd/codaw render  testdata/basic/project.toml old.wav   # existing engine
go run ./cmd/codaw render2 testdata/basic/project.toml new.wav   # new mixer

Compare with a tolerance, not byte equality — different summing order changes the last bits of a float. Peak and RMS per second, within ~0.01 dB, is a good check.

Expect the reverb/EQ tracks to differ (no FX yet — that's issue 3). Test gain, pan, routing and timing on dry tracks first.

Why this solution

  • You can read the whole signal path in one screen. That was the explicit goal. No interfaces, no indirection, no framework.
  • It mirrors the TOML schema exactly. project.toml describes tracks, buses and a master; so does the loop. There's no mental translation step.
  • It's testable with no hardware. Build a Project in a test, call MixBlock, assert on floats. None of that is possible today.
  • Zero allocation is easy to reach and easy to prove — everything is preallocated in New, and testing.AllocsPerRun pins it.

Alternatives considered

A streamer tree — every node implements Read(buf []float32), master pulls buses, buses pull tracks (this is how gopxl/beep works). Genuinely elegant, and render-equals-playback falls out structurally. Rejected because the same guarantee is available here by just having both callers call MixBlock, and the flat version has no end-of-stream/short-read subtleties and no allocation hiding behind interface calls. Elegance that costs understanding is a bad trade for this project.

A generic node graph with ports and a topological sort. That's what miniaudio was doing — it handles sends, sidechains and feedback. Rejected: the TOML schema can't express a send, so it would be machinery for a feature that doesn't exist. If sends ever land, this loop grows into it.

Keep per-parameter events driving a mutable graph. Rejected. The mixer reads the immutable snapshot from Store.Get() every block, so it simply sees new values — nothing needs to be told what changed. This is why most of state/events.go becomes dead (issue 5).

Sample-accurate automation (ramping within a block rather than once per block). Correct, and audibly better on fast fades. Deferred — per-block is already far better than the 20 ms ticker, and this can be added later inside MixBlock without changing anything around it.

Done when

  • mix.MixBlock() renders clips, tracks, buses and master with gain, pan, mute/solo
  • Automation sampled once per block from the frame clock — no time.Ticker anywhere
  • 5 ms fade at every clip edge
  • Constant-power pan law
  • render2 output matches render on dry tracks within tolerance
  • testing.AllocsPerRun(100, mixBlock) == 0
  • Unit tests: silence in → silence out; a −6 dB track is half amplitude; hard-left pan puts nothing in the right channel; a looped clip wraps at the right frame

Not in this issue

Effects (issue 3), the device (issue 4), seeking beyond what render needs. Keep codaw play on the old engine — nothing is deleted until issue 4.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions