diff --git a/README.md b/README.md index 20d2e0d..2c8801a 100644 --- a/README.md +++ b/README.md @@ -29,7 +29,7 @@ async function main() { const modelPath = await Model.download('quail-vf-2.2-s-16khz', './models') const model = Model.fromFile(modelPath) - // The model's own settings give the lowest delay + // Use the model's optimal configuration for the lowest delay const sampleRate = model.getOptimalSampleRate() const blockSize = model.getOptimalBlockSize(sampleRate) @@ -45,20 +45,22 @@ main() ``` Processing is mono. For multichannel audio, mix down to mono or use one processor per -channel. +channel. Initialize each instance before passing audio. Blocks must contain exactly +`blockSize` samples. Set the third `initialize` argument, `variableBlockSize`, to `true` +to allow shorter blocks; longer blocks are always rejected. This also applies to VAD +processing and analyzer buffering. -The snippets below are fragments, and each one leaves out that surrounding `async function` -for brevity. Keep it: `await` at the top level of a CommonJS file is a syntax error, and -several calls here are asynchronous. +The following snippets assume an enclosing `async` function. CommonJS files do not support +top-level `await`. Runnable scripts for enhancement, VAD, analysis and whole-file processing, synchronous and async, are in [`examples/`](examples). ## Models -Available models and their ids are listed at -[artifacts.ai-coustics.io](https://artifacts.ai-coustics.io). Each class accepts exactly one -family of models and throws for the rest: +Available models and their IDs are listed at +[artifacts.ai-coustics.io](https://artifacts.ai-coustics.io). Each class accepts the model types +listed below and rejects other types: | Class | Accepted models | | ----------------------------- | ------------------- | @@ -66,9 +68,9 @@ family of models and throws for the rest: | `Vad`, `VadAsync` | dedicated VAD | | `Analyzer` | analysis | -`Model.download` resolves to the model's path and runs off the event loop. Model files are -memory-mapped rather than read into memory, so keep the file in place while anything created -from it is alive. +`Model.download` runs on Node's libuv thread pool and returns a promise for the model's path. +`Model.fromFile` memory-maps the file. Do not modify or delete it while the model or any +instance created from it is still alive. ## Enhancement @@ -84,40 +86,38 @@ context.setParameter(ProcessorParameter.Bypass, 0) console.log(context.getParameter(ProcessorParameter.EnhancementLevel)) -// Samples of delay the processor adds, for lining the output up with other streams +// Get the audio delay in samples to align the output with other streams console.log(context.getAudioDelay()) // Clear internal state on a stream discontinuity or seek context.reset() ``` -## Off the main thread +## Async processing -`ProcessorAsync` and `VadAsync` do the same work as `Processor` and `Vad`, but each call -returns a promise and runs on a worker thread, so the event loop stays responsive. Use them -when other work shares that loop, as in a server handling live streams alongside its sockets -and HTTP. The synchronous classes are the pick on a dedicated audio thread or in a batch -script, -where nothing else needs the loop and a promise per block would only add overhead. +`ProcessorAsync` and `VadAsync` run initialization and processing on Node's libuv thread pool. +Await these calls to keep the event loop available for other work, such as network requests. +Construction and disposal remain synchronous. + +Use `Processor` or `Vad` on a dedicated worker thread, or in a batch script that can block +the calling thread. These classes process each block without a promise or input copy. ```javascript const { Model, ProcessorAsync } = require('@ai-coustics/aic-sdk') const processor = await new ProcessorAsync(model, process.env.AIC_SDK_LICENSE).withConfig(sampleRate, blockSize) -// Unlike the synchronous class, this does not write into the caller's array. It resolves -// to the enhanced samples, so a streaming loop reuses the same variable -let audio = new Float32Array(blockSize) -for (;;) { - audio = await processor.process(audio) -} +// The input remains unmodified; the promise resolves to the enhanced samples. +const audio = new Float32Array(blockSize) +const enhanced = await processor.process(audio) ``` -The input is copied before the work is queued, so it stays valid and untouched while the -promise is pending. `VadAsync.process` hands the block back unmodified in the same way. +Both async classes copy the input before queuing work and return a new `Float32Array`. +`ProcessorAsync.process` returns enhanced samples; `VadAsync.process` returns the original +samples. The caller's input remains unmodified. -Context handles are awaited but their methods stay synchronous, so a prediction or parameter -can still be read from inside an audio callback: +Context creation is asynchronous. The returned context's methods are synchronous and can +be used while processing runs: ```javascript const context = await processor.getContext() @@ -126,9 +126,9 @@ context.setParameter(ProcessorParameter.EnhancementLevel, 0.8) ### Running several streams -One instance handles one stream. Do not start a second `process` on the same instance before -the first resolves: worker threads complete out of order, which would desync the stream. Give -each stream its own instance instead. +Use one instance per stream and await each operation before submitting the next. Concurrent +calls on one instance are not guaranteed to execute in submission order. Separate instances +can process streams concurrently. ```javascript const processors = await Promise.all( @@ -146,9 +146,7 @@ Node starts: UV_THREADPOOL_SIZE=16 node server.js ``` -The core SDK's own `AIC_NUM_THREADS` has no effect here: it sizes a thread pool these -bindings deliberately do not use, so that all audio work stays on the pool Node already -manages. +`AIC_NUM_THREADS` does not apply to these bindings; async audio work uses libuv. ## Voice activity detection @@ -182,7 +180,8 @@ When enhancement and detection run together, feed the VAD the **original** input processor's output. Enhancement is designed to change the signal, so detecting on its output means running the VAD on audio that no longer matches what its model expects, and it stacks the processor's delay on top of the prediction delay. Because `vad.process` leaves its input -untouched, running both on the same block is enough: +untouched, run it before enhancement. Configure both instances with the same sample rate and +block size: ```javascript vad.process(block) // reads the block @@ -196,13 +195,13 @@ context.getAudioDelay() // enhanced audio lags the input by this many samples vadContext.getPredictionDelay() // the VAD decision lags the same input by this many ``` -The prediction delay is not applied to the audio; use it to line speech decisions up with -the audio timeline. +The prediction delay is not applied to the audio. Use it to align speech decisions with +the input audio. ## Analysis -Analysis models score audio quality. Buffering is cheap enough for the audio path; -`analyze` runs the model and is not. +Analysis models score audio quality. Collect audio with `buffer`, then run the model with +`analyzeAsync` or `analyze`. ```javascript const { Model, Analyzer } = require('@ai-coustics/aic-sdk') @@ -220,19 +219,21 @@ const result = await analyzer.analyzeAsync() console.log(result.riskScore, result.noise, result.speakerReverb) ``` -`buffer` is cheap enough for the audio path; running the model is not, so it is a separate -call with two forms. `analyzeAsync` is the one to reach for in a server: analysis is -occasional, so the promise costs nothing next to the model, and the event loop stays free. -`analyze` blocks and suits a CLI or a worker thread. +`buffer` is synchronous and does not acquire the analyzer lock. Audio collection can +continue while `analyzeAsync` runs on a worker thread. Use `analyze` when blocking the +calling thread is acceptable, such as in a CLI or dedicated worker. -Only that one call moves off-thread, which is why there is no `AnalyzerAsync` class to match -`ProcessorAsync`. `buffer` stays synchronous, takes no lock, and can be called while an -analysis is still running. The SDK guarantees the collector and analyzer halves are safe to -use concurrently. The analyzer's other methods do wait for a pending analysis to finish. +Analysis uses a fixed duration of audio determined by the model. Older samples are discarded +as new audio arrives. If insufficient audio has been collected, analysis pads the remaining +input with silence. -Every score runs 0.0 - 1.0. Except `speakerLoudness`, lower means less problematic audio. -`riskScore` is the headline number: how likely this audio is to break downstream models such -as speech-to-text, VAD or turn-taking. +Calls to `analyze`, `reset`, `updateBearerToken`, `terminateSession` and `dispose` acquire the +analyzer lock and may block while async analysis is running. `initialize` only configures +the collector. + +All scores range from 0.0 to 1.0. For every field except `speakerLoudness`, lower values +indicate less problematic audio. `riskScore` predicts the likelihood of failure in downstream +models such as speech-to-text, VAD or turn-taking. ## Telemetry @@ -247,10 +248,11 @@ const processor = new Processor(model, licenseKey, { }) ``` -A session is closed when its object is garbage collected. Because GC timing is not -guaranteed, every processor, VAD and analyzer also exposes `terminateSession()` for -lifecycle events; afterwards the object can no longer process audio. On `ProcessorAsync` -and `VadAsync` it returns a promise, since it may block. +A telemetry session ends when its native object is destroyed. Use `terminateSession()` to +request termination at a specific lifecycle event. Once termination is handled, processors +and VADs can no longer process audio, and analyzers can no longer analyze buffered audio. +On `ProcessorAsync` and `VadAsync`, termination runs on a libuv worker thread and returns +a promise because it may block. If your license key is a JWT, refresh it in place instead of rebuilding the object: @@ -261,13 +263,12 @@ context.updateBearerToken(renewedJwt) ## Memory management `Model`, `Processor`, `ProcessorAsync`, `Vad`, `VadAsync` and `Analyzer` hold large native -allocations behind small JavaScript objects. The binding reports each object's native -footprint to V8 (`napi_adjust_external_memory`), so the garbage collector applies the right -amount of pressure and reclaims dropped instances promptly instead of letting native memory -grow unbounded. +allocations behind small JavaScript objects. The binding reports estimated native memory +usage to V8 so the garbage collector can account for these allocations. This influences +collection frequency but does not guarantee when an object will be released. -For deterministic cleanup, every one of these classes also exposes `dispose()`, which -destroys the native object immediately instead of waiting for garbage collection: +Use `dispose()` to release native resources at a specific point instead of waiting for +garbage collection: ```javascript const processor = new Processor(model, licenseKey) @@ -279,10 +280,14 @@ try { } ``` -After `dispose()`, every method on the object throws; calling `dispose()` again does -nothing. On the async classes it blocks until in-flight work on the libuv pool finishes. +After `dispose()`, all methods except `dispose()` fail; repeated disposal has no effect. +Async methods reject their promises. Disposal blocks the calling thread if a worker holds +the native object's lock. Queued work that acquires the lock after disposal fails. + +Disposing a model releases its reference to the model data. Processors, VADs and analyzers +created from it retain their own references and remain usable. -Two things to know about cleanup timing: +Cleanup timing also affects native memory usage: - Native cleanup runs on the event loop when the object is finalized, not synchronously at garbage collection. Finalizers run on event-loop turns, so batches that create many of @@ -291,9 +296,9 @@ Two things to know about cleanup timing: `setTimeout` or `setImmediate` are. A tight synchronous loop that creates thousands of objects accumulates their native memory for the duration of the loop; create these objects per unit of work behind real async boundaries, or reuse a single instance. -- RSS reflects the peak of simultaneously live (or not-yet-finalized) instances: the - allocator reuses freed native memory rather than returning it to the OS, so a burst of N - concurrent instances costs about N x their footprint even after they are dropped. +- RSS may remain elevated after disposal because the allocator can retain freed memory + for reuse. Peak memory use depends on the number of concurrently live instances, + including objects awaiting finalization. ## Development diff --git a/__test__/common.ts b/__test__/common.ts index 8f5c2c3..72b7134 100644 --- a/__test__/common.ts +++ b/__test__/common.ts @@ -27,9 +27,9 @@ export function modelPath(kind: ModelKind): string { } /** - * The license key the SDK needs to construct anything. + * Returns the SDK license key used by the tests. * - * Throws up front, so a missing key is not reported later as a license error. + * Throws if AIC_SDK_LICENSE is missing. */ export function licenseKey(): string { const key = process.env.AIC_SDK_LICENSE diff --git a/__test__/dispose.spec.ts b/__test__/dispose.spec.ts index 350aa33..3c80f2b 100644 --- a/__test__/dispose.spec.ts +++ b/__test__/dispose.spec.ts @@ -190,8 +190,7 @@ test('analyzer dispose waits for an in-flight analyzeAsync to finish', async (t) analyzer.dispose() const blockedMs = elapsedMs(disposeStarted) - // dispose() waited for the lock instead of pulling the analyzer out from under the - // worker, so the analysis produced a real result. + // The worker completed analysis before disposal acquired the lock. const result = await inFlight t.is(typeof result.riskScore, 'number', 'the in-flight analysis must complete, not be cancelled') @@ -210,8 +209,8 @@ test('a withConfig handle outlives its original handle being collected', async ( const sampleRate = model.getOptimalSampleRate() const blockSize = model.getOptimalBlockSize(sampleRate) - // The chaining form leaves the constructor's handle as garbage. Its finalizer must - // withdraw only its own footprint report, not destroy the shared native processor. + // The constructor's temporary handle may be finalized after `withConfig` resolves. + // Its finalizer must preserve the shared processor and its single memory report. const processor = await new ProcessorAsync(model, licenseKey()).withConfig(sampleRate, blockSize) // Push V8 towards a GC so the dropped handle's finalizer runs. diff --git a/__test__/index.spec.ts b/__test__/index.spec.ts index 6caf151..8b8be2b 100644 --- a/__test__/index.spec.ts +++ b/__test__/index.spec.ts @@ -44,8 +44,8 @@ test('sdk expects the model version the fixtures are published under', (t) => { test('model exposes its id and optimal settings', (t) => { const model = enhancementModel() - // The model file's own id carries the build hash and version, e.g. - // `quail-vf-2.2-s-16khz-gf70x7zf-v14`, so it extends the manifest id used to download it. + // The model file's own ID carries the build hash and version, e.g. + // `quail-vf-2.2-s-16khz-gf70x7zf-v14`, so it extends the manifest ID used to download it. t.true( model.getId().startsWith(TEST_MODELS.enhancement.id), `expected ${model.getId()} to start with ${TEST_MODELS.enhancement.id}`, @@ -175,9 +175,8 @@ test('async processor matches the sync processor sample for sample', async (t) = model.getOptimalBlockSize(model.getOptimalSampleRate()), ) - // Same model, settings and input, and both start from a fresh state, so moving the work - // to a worker thread and copying the block across must not change the output. Four - // blocks, so a drift in the processor's internal state would show up too. + // Compare four blocks from fresh processors with identical input and settings. + // Async scheduling and buffer copies must preserve output and state transitions. for (let block = 0; block < 4; block += 1) { const audio = ramp(blockSize) const expected = audio.slice() @@ -216,8 +215,7 @@ test('async processors on separate instances can run concurrently', async (t) => const sampleRate = model.getOptimalSampleRate() const blockSize = model.getOptimalBlockSize(sampleRate) - // Parallelism is across instances, never within one: libuv completes work items out of - // order, so overlapping calls on a single instance would desync the stream. + // Use independent processors for concurrency; calls within each stream must be ordered. const processors = await Promise.all( Array.from({ length: 4 }, () => new ProcessorAsync(model, licenseKey()).withConfig(sampleRate, blockSize)), ) @@ -408,10 +406,8 @@ test('audio can be buffered while an analysis is in flight', async (t) => { t.is(typeof result.riskScore, 'number') }) -// 'async processor rejects instead of throwing' covers async errors arriving as a -// rejection, because processing before initialize is a certain error. The analyzer has no -// equally certain one: nobody has checked whether analyzing before initialize errors or -// returns silence-padded scores, so this file asserts nothing about it. +// Async rejection behavior is covered by the processor test, which processes before +// initialization. No equivalent analyzer failure is assumed here. test('each class rejects the wrong model type', (t) => { const key = licenseKey() diff --git a/__test__/models.ts b/__test__/models.ts index b5208dd..0cb8c70 100644 --- a/__test__/models.ts +++ b/__test__/models.ts @@ -15,7 +15,7 @@ export interface TestModel { * Model fixtures the end-to-end tests run against. * * `scripts/fetch-test-models.mjs` resolves these through `Model.download()`, which - * re-fetches the manifest and pulls the newest compatible model version. The model file + * fetches the manifest and selects the latest compatible model version. The model file * format version is tied to the SDK version, so `getCompatibleModelVersion()` reports * the version the built addon expects, and the `sdk expects the model version the * fixtures are published under` test asserts the fixtures still match it. diff --git a/benchmark/bench.ts b/benchmark/bench.ts index 570d27f..8461bbb 100644 --- a/benchmark/bench.ts +++ b/benchmark/bench.ts @@ -24,8 +24,7 @@ const blockSize = model.getOptimalBlockSize(sampleRate) const processor = new Processor(model, licenseKey) processor.initialize(sampleRate, blockSize) -// How many streams the concurrent case runs at once. Each stream gets its own processor, -// since overlapping calls on one instance would desync it. +// Number of concurrent streams. Each needs an independent processor to preserve block order. const concurrency = Number(process.env.AIC_BENCH_CONCURRENCY ?? 4) const asyncProcessor = await new ProcessorAsync(model, licenseKey).withConfig(sampleRate, blockSize) @@ -33,27 +32,21 @@ const asyncProcessors = await Promise.all( Array.from({ length: concurrency }, () => new ProcessorAsync(model, licenseKey).withConfig(sampleRate, blockSize)), ) -// One block of speech-like content, reused so the benchmark measures processing rather -// than buffer allocation. Nothing writes into it, so every task starts from the same -// reference signal. +// Reuse a reference signal to exclude input allocation from the measurement. const audio = Float32Array.from({ length: blockSize }, (_, i) => Math.sin(i / 10) * 0.5) -// The synchronous call enhances in place, so it gets its own scratch buffer, refilled from -// `audio` before every iteration. Enhancing one shared buffer would instead feed each -// iteration the previous one's output, and would leave the async tasks below measuring -// audio that had already been enhanced hundreds of thousands of times. +// Refill the synchronous processor's scratch buffer before each iteration. +// This prevents enhanced output from becoming the next iteration's input. const syncAudio = new Float32Array(blockSize) -// The async calls resolve to a fresh array each time instead of writing in place, so each -// stream keeps its own block to hand back in. +// Keep a separate result buffer for each async stream. const asyncAudio = asyncProcessors.map(() => audio.slice()) const syncTask = `sync: 1 block` const asyncTask = `async: 1 block` const concurrentTask = `async: ${concurrency} blocks concurrently` -// Blocks of audio each iteration gets through, so the real-time factor below can compare -// the one-block cases against the concurrent one on equal terms. +// Track blocks per iteration to normalize throughput across benchmark cases. const blocksPerIteration = new Map([ [syncTask, 1], [asyncTask, 1], @@ -71,8 +64,7 @@ bench.add( { beforeEach: () => syncAudio.set(audio) }, ) -// Same work as above on a worker thread, so the gap against `sync` is the cost of the -// promise plus the copy in and out. +// Measure async scheduling, promise and input-copy overhead relative to synchronous processing. bench.add(asyncTask, async () => { await asyncProcessor.process(audio) }) @@ -83,10 +75,8 @@ bench.add(concurrentTask, async () => { await Promise.all(asyncProcessors.map((instance, stream) => instance.process(asyncAudio[stream]))) }) -// Analysis, if an analysis model was supplied. Measured separately from enhancement: -// `analyze` is an occasional call over a span of audio, not a per-block one, and its cost -// decides whether `analyzeAsync` is needed. Anything in the tens of milliseconds is too -// long to sit on the event loop. +// Benchmark analysis separately when an analysis model is supplied. +// Analysis processes a fixed duration of buffered audio per call. const analysisModelPath = process.env.AIC_SDK_ANALYSIS_MODEL if (analysisModelPath) { const analysisModel = Model.fromFile(analysisModelPath) @@ -107,8 +97,7 @@ if (analysisModelPath) { analyzer.analyze() }) - // The gap against the blocking call is the promise plus the thread hop. Against a model - // this size it should be lost in the noise. + // Measure async scheduling overhead relative to synchronous analysis. bench.add('analysis: analyzeAsync', async () => { await analyzer.analyzeAsync() }) diff --git a/examples/README.md b/examples/README.md index bd6ec68..affd5d2 100644 --- a/examples/README.md +++ b/examples/README.md @@ -2,19 +2,17 @@ Runnable scripts for each part of the SDK. -| Example | Shows | -| ---------------------- | --------------------------------------------------------- | -| `enhancement.js` | Speech enhancement, in place on the calling thread | -| `enhancement-async.js` | The same on a worker thread, then several streams at once | -| `vad.js` | Voice activity detection and its parameters | -| `vad-async.js` | Async detection, and detection combined with enhancement | -| `analysis.js` | Audio quality scoring, blocking and on a worker thread | -| `file-processing.js` | A WAV file end to end, with delay compensation | - -Analysis needs no separate async script: only `analyzeAsync` moves off the calling thread, so -both forms live in `analysis.js`. File processing has no async counterpart because it does not need -one: a batch job has no event loop to keep free, and the sync API is the simpler tool. For -parallel batch work, run several processes. +| Example | Shows | +| ---------------------- | -------------------------------------------------------- | +| `enhancement.js` | Speech enhancement, in place on the calling thread | +| `enhancement-async.js` | Async speech enhancement and concurrent streams | +| `vad.js` | Voice activity detection and its parameters | +| `vad-async.js` | Async detection, and detection combined with enhancement | +| `analysis.js` | Audio quality scoring, blocking and on a worker thread | +| `file-processing.js` | WAV file enhancement with delay compensation | + +`analysis.js` demonstrates both `analyze` and `analyzeAsync`. The file processing example +uses synchronous processing. For parallel batch processing, run several processes. ## Setup @@ -54,26 +52,24 @@ node examples/file-processing.js --input speech.wav node examples/file-processing.js --input speech.wav --output enhanced.wav --enhancement 0.7 ``` -`--model` tries a different model and `--help` lists every option. A model only enhances up +`--model` selects the model and `--help` lists all options. A model only enhances up to its own Nyquist limit, so pair a 48 kHz source with a 48 kHz model such as `rook-l-48khz`. Browse the catalogue at [artifacts.ai-coustics.io](https://artifacts.ai-coustics.io). ## Choosing between sync and async -`Processor` and `Vad` do their work on the thread that calls them. That is what you want on a -dedicated audio thread, where a promise per block would only add overhead. +`Processor` and `Vad` run on the calling thread. Use them in a dedicated worker or a batch +script where blocking is acceptable. -`ProcessorAsync` and `VadAsync` run on Node's libuv thread pool, so the event loop stays -responsive. That matters when something else is waiting on it, as in a server handling live -streams alongside its sockets and HTTP. It buys a batch script nothing. Two differences to -keep in mind: +`ProcessorAsync` and `VadAsync` run processing on Node's libuv thread pool, keeping the event +loop available for other work. Their constructors and `dispose()` methods are synchronous. -- `process` does not write into the array it is given. It copies the input, so that array - stays valid while the promise is pending, and resolves to the samples instead. -- One instance handles one stream, and calls on it must not overlap: worker threads finish - out of order, which would desync the stream. Parallelism comes from running several - instances, as the end of `enhancement-async.js` shows. +- `process` copies the input and returns a promise for a new array. The input remains + unmodified. Enhancement returns enhanced samples; VAD returns the original samples. +- Await each operation before submitting the next on the same instance. Calls are not + guaranteed to execute in submission order. Use one instance per stream to process + multiple streams concurrently. -The pool is four threads by default and is shared with `fs`, `dns` and `crypto`. Raise -`UV_THREADPOOL_SIZE` before Node starts to run more streams in parallel. The core SDK's -`AIC_NUM_THREADS` has no effect on these bindings. +The pool defaults to four threads and is shared with filesystem, DNS and crypto work. Set +`UV_THREADPOOL_SIZE` before starting Node to change its size. `AIC_NUM_THREADS` does not apply +to these bindings. diff --git a/examples/analysis.js b/examples/analysis.js index 8b9f0a9..3845364 100644 --- a/examples/analysis.js +++ b/examples/analysis.js @@ -1,11 +1,7 @@ -// Audio quality analysis. +// Audio quality analysis on the calling thread and on a libuv worker. // -// Buffering and analysis are separate calls: `buffer` is cheap enough for the audio path, -// while running the model is not. Analysis comes in two forms, both shown below: -// `analyzeAsync` on a worker thread and `analyze` on the calling thread. -// -// There is no separate `AnalyzerAsync` class because only one call moves off-thread. -// `buffer` stays synchronous, and can be called while an analysis runs. +// `buffer` collects audio synchronously and can continue during `analyzeAsync`. +// Analysis is computationally expensive; avoid running `analyze` in audio processing callbacks. const { Analyzer, Model, getVersion } = require('..') @@ -44,8 +40,7 @@ async function main() { } console.log('Buffered 100 blocks') - // The heavy call, on a worker thread. Use this in a server: the event loop stays free - // while the model runs, and the tick counter below shows it. + // Run analysis on a worker and count event-loop timer callbacks while it runs. let ticks = 0 const ticker = setInterval(() => { ticks += 1 @@ -57,9 +52,8 @@ async function main() { clearInterval(ticker) - // Every score runs 0.0 - 1.0. Except `speakerLoudness`, lower means less problematic - // audio. `riskScore` is the headline number: how likely this audio is to break downstream - // models such as speech-to-text, VAD or turn-taking. + // Scores range from 0.0 to 1.0. Lower values indicate fewer problems, except for + // speakerLoudness. riskScore predicts failure in downstream speech models. console.log('\nAnalysis:') for (const [name, score] of Object.entries(result)) { console.log(` ${name.padEnd(20)} ${score.toFixed(4)}`) @@ -71,8 +65,7 @@ async function main() { const syncElapsed = performance.now() - syncStart console.log(`\nanalyzeAsync: ${asyncElapsed.toFixed(1)} ms, analyze: ${syncElapsed.toFixed(1)} ms`) - // The blocking call would have held the event loop for its whole duration; the async one - // does not, which is what these ticks show. + // Report how many timer callbacks ran during async analysis. console.log(`Timer fired ${ticks} times during analyzeAsync`) // `buffer` drives the collector, not the analyzer, so audio can keep arriving while an diff --git a/examples/audio-file.js b/examples/audio-file.js index a1cd6c8..706013d 100644 --- a/examples/audio-file.js +++ b/examples/audio-file.js @@ -1,10 +1,7 @@ // WAV reading and writing for file-processing.js. // -// Nothing here is SDK-specific; it is kept in its own file so the example stays about the -// SDK. `wavefile` is a devDependency of this repo; install it alongside the SDK to run -// that example: -// -// npm install wavefile +// Requires `wavefile`, included in this repository's devDependencies. +// For use outside this repository: npm install wavefile const fs = require('node:fs') @@ -19,11 +16,10 @@ const { WaveFile } = require('wavefile') function readWav(path) { const wav = new WaveFile(fs.readFileSync(path)) - // Normalizes 16-bit PCM, 24-bit, float and the rest to the float samples the SDK wants. + // Convert PCM and floating-point WAV samples to the floating-point format used by the SDK. wav.toBitDepth('32f') - // `getSamples` is declared as returning Float64Array whatever container it is handed, so - // the Float32Array it actually produces has to be restated for the type checker. + // The wavefile type declaration does not reflect the requested Float32Array output. const samples = /** @type {Float32Array[] | Float32Array} */ ( /** @type {unknown} */ (wav.getSamples(false, Float32Array)) ) @@ -32,8 +28,7 @@ function readWav(path) { const { sampleRate } = /** @type {{ sampleRate: number }} */ (wav.fmt) return { - // De-interleaved multichannel arrives as an array of channels, but mono arrives as one - // flat array, so it is wrapped to give callers a single shape to handle. + // Return an array of channels for both mono and multichannel input. channels: Array.isArray(samples) ? samples : [samples], sampleRate, } @@ -49,8 +44,7 @@ function readWav(path) { function writeWav(path, channels, sampleRate) { const wav = new WaveFile() - // `fromScratch` wants an array of channels for multichannel but a flat array for mono, - // the mirror image of what `readWav` normalizes away. + // `fromScratch` expects a flat array for mono and an array of channels otherwise. wav.fromScratch(channels.length, sampleRate, '32f', channels.length === 1 ? channels[0] : channels) fs.writeFileSync(path, wav.toBuffer()) diff --git a/examples/enhancement-async.js b/examples/enhancement-async.js index 5410fec..b18c5e0 100644 --- a/examples/enhancement-async.js +++ b/examples/enhancement-async.js @@ -1,13 +1,8 @@ -// Speech enhancement off the main thread. +// Speech enhancement on Node's libuv thread pool. // -// `ProcessorAsync` does the same work as `Processor`, but on a worker thread, so the event -// loop stays free. Two differences to watch for: -// -// * `process` does not write into the caller's array. It copies the input, so the array -// stays valid while the promise is pending, and resolves to the enhanced samples. -// * `getContext()` is awaited, but the handle it resolves to is fully synchronous. -// -// It ends by running several streams at once, the main reason to use the async API. +// `process` copies its input and returns a promise for the enhanced samples. +// `getContext` returns a promise; the context's methods are synchronous. +// The final example processes several independent streams concurrently. const os = require('node:os') @@ -37,26 +32,24 @@ async function main() { const blockSize = model.getOptimalBlockSize(sampleRate) console.log(`Audio format: ${blockSize} samples @ ${sampleRate} Hz`) - // `withConfig` initializes and resolves to a handle, so construction and setup chain into - // one await. `new ProcessorAsync(...)` plus `await processor.initialize(...)` is equivalent. + // Create and initialize the processor. Calling `initialize` on a constructed + // instance is equivalent. const processor = await new ProcessorAsync(model, licenseKey).withConfig(sampleRate, blockSize) const context = await processor.getContext() console.log('Audio delay:', context.getAudioDelay(), 'samples') - // Synchronous even on the async class, so parameters can be changed from anywhere, - // including from inside an audio callback. + // Context methods are synchronous and can be used during processing. context.setParameter(ProcessorParameter.EnhancementLevel, 0.7) console.log('Enhancement level:', context.getParameter(ProcessorParameter.EnhancementLevel)) - // A timer to prove the event loop is not blocked while the model runs. + // Count event-loop timer callbacks during processing. let ticks = 0 const ticker = setInterval(() => { ticks += 1 }, 1) - // The steady-state streaming loop. `process` resolves to a new array, so the block is - // reassigned instead of mutated, and one variable can carry the stream. + // Await each block in sequence. The result is a new array; the input is unmodified. let audio = Float32Array.from({ length: blockSize }, () => (Math.random() - 0.5) * 0.2) console.log('Before:', audio.slice(0, 4)) @@ -73,15 +66,11 @@ async function main() { const audioMs = (BLOCKS * blockSize * 1000) / sampleRate console.log(`\nProcessed ${BLOCKS} blocks (${audioMs.toFixed(0)} ms of audio) in ${elapsed.toFixed(0)} ms`) console.log(`Real-time factor: ${(audioMs / elapsed).toFixed(1)}x`) - // Zero here would mean the work had blocked the event loop. + // Report how many timer callbacks ran during processing. console.log(`Timer fired ${ticks} times while processing, so the event loop stayed responsive`) - // Several streams at once. - // - // One instance handles one stream, and calls on it must not overlap: worker threads - // finish out of order, so a second `process` before the first resolves would desync the - // stream. Concurrency comes from running several instances side by side, one per stream, - // as a server would do per connection. + // Process independent streams concurrently, with one processor per stream. + // Await calls within each stream to preserve block order. console.log(`\nEnhancing ${STREAMS} streams concurrently`) const processors = await Promise.all( @@ -100,8 +89,7 @@ async function main() { ) const concurrentElapsed = performance.now() - concurrentStart - // Bounded by the libuv pool, four threads unless UV_THREADPOOL_SIZE says otherwise, so - // raising it lets more streams overlap. + // Concurrency is limited by the libuv pool size, configured with UV_THREADPOOL_SIZE. console.log(`${STREAMS} x ${BLOCKS} blocks in ${concurrentElapsed.toFixed(0)} ms`) console.log(`Aggregate: ${((STREAMS * audioMs) / concurrentElapsed).toFixed(1)}x real time`) console.log(`libuv pool: ${process.env.UV_THREADPOOL_SIZE ?? '4 (default)'}, ${os.cpus().length} logical CPUs`) diff --git a/examples/enhancement.js b/examples/enhancement.js index 927cb2a..dd2d0a7 100644 --- a/examples/enhancement.js +++ b/examples/enhancement.js @@ -1,8 +1,7 @@ -// Speech enhancement on the main thread. +// Speech enhancement on the calling thread. // -// The synchronous API: `process` enhances a block in place and returns nothing. Use this -// when you are already on a dedicated audio thread. For a server or batch job, see -// enhancement-async.js. +// `process` modifies a mono block in place. Use this API in a dedicated worker or +// batch script. See enhancement-async.js for processing through the libuv thread pool. const { Model, Processor, ProcessorParameter, getCompatibleModelVersion, getVersion } = require('..') @@ -29,7 +28,7 @@ async function main() { const model = Model.fromFile(modelPath) console.log('Model id:', model.getId()) - // The model's own settings give the lowest delay. + // Use the model's optimal configuration for the lowest delay. const sampleRate = model.getOptimalSampleRate() const blockSize = model.getOptimalBlockSize(sampleRate) console.log(`Audio format: ${blockSize} samples @ ${sampleRate} Hz`) @@ -45,7 +44,7 @@ async function main() { context.setParameter(ProcessorParameter.EnhancementLevel, 0.7) console.log('Enhancement level:', context.getParameter(ProcessorParameter.EnhancementLevel)) - // Stand-in for real audio: noise at a low level. Feed this real speech to hear anything. + // Generate low-level noise as sample input. Replace this with audio from your source. const audio = Float32Array.from({ length: blockSize }, () => (Math.random() - 0.5) * 0.2) console.log('Before:', audio.slice(0, 4)) diff --git a/examples/file-processing.js b/examples/file-processing.js index ad2c557..8d86bad 100644 --- a/examples/file-processing.js +++ b/examples/file-processing.js @@ -1,7 +1,4 @@ -// Enhances a WAV file end to end. -// -// Shows the two things a file tool needs beyond the basic loop: one processor per channel, -// and compensating for the processor's delay so the output lines up with the input. +// Enhances a WAV file using one processor per channel and compensates for audio delay. // // Usage: // node examples/file-processing.js --input speech.wav [--output enhanced.wav] @@ -30,8 +27,8 @@ model such as rook-l-48khz. Browse models at https://artifacts.ai-coustics.io` * @typedef {object} Options * @property {string} [input] WAV file to enhance * @property {string} [output] Where to write the result - * @property {string} model Model id - * @property {number} enhancement Enhancement level, 0.0 - 1.0 + * @property {string} model Model ID + * @property {number} enhancement Enhancement level, 0.0 to 1.0 * @property {boolean} [help] Whether usage was requested */ @@ -67,7 +64,7 @@ function parseArgs(argv) { options.help = true break default: - // A lone path with no flag is taken as the input, so the common case needs no flags. + // Accept a positional input path. if (!arg.startsWith('-') && !options.input) { options.input = arg } else { @@ -116,8 +113,8 @@ async function main() { const model = Model.fromFile(await Model.download(options.model, MODEL_DIR)) console.log('Model id:', model.getId()) - // The file's rate drives the format, not the model's, since the SDK resamples - // internally. The block size that avoids extra buffering depends on that rate. + // Use the file's sample rate; the SDK resamples internally as needed. + // The optimal block size depends on this rate. const blockSize = model.getOptimalBlockSize(sampleRate) // Processing is mono, so each channel gets its own processor and its own internal state. @@ -131,9 +128,8 @@ async function main() { return { processor, context } }) - // Enhanced audio comes out this many samples late. Feeding that many extra samples of - // silence flushes the tail, and skipping that many samples of output realigns the result - // with the input. Without this the file would be shifted and truncated. + // Append silence to flush the delayed output, then skip the initial delay samples + // to align the enhanced file with the input. const delay = processors[0].context.getAudioDelay() console.log(`Block size: ${blockSize}, enhancement level: ${options.enhancement}, delay: ${delay} samples`) @@ -146,8 +142,7 @@ async function main() { padded.set(channel) for (let offset = 0; offset < paddedLength; offset += blockSize) { - // A view onto `padded`, not a copy, and `process` enhances in place, so the result - // is written straight back into the output buffer. + // `subarray` shares the output buffer, so in-place processing writes directly into it. processors[index].processor.process(padded.subarray(offset, offset + blockSize)) } diff --git a/examples/vad-async.js b/examples/vad-async.js index f689124..0533681 100644 --- a/examples/vad-async.js +++ b/examples/vad-async.js @@ -1,8 +1,7 @@ -// Voice activity detection off the main thread. +// Voice activity detection on Node's libuv thread pool. // -// `VadAsync` mirrors `Vad` on a worker thread. `process` copies the block in and hands it -// straight back, so the same block can go to the VAD and then on to a processor, as in the -// combined pattern at the end of this file. +// `process` returns a promise for a copy of the original samples. The final example +// passes this audio to an enhancement processor after detection. const { Model, ProcessorAsync, VadAsync, VadParameter, getVersion } = require('..') @@ -29,8 +28,7 @@ async function main() { const vad = await new VadAsync(model, licenseKey).withConfig(sampleRate, blockSize) - // Awaited once; every method on the handle itself is synchronous, so predictions can be - // read from anywhere, including from inside an audio callback. + // Await context creation. Prediction and parameter methods are synchronous. const context = await vad.getContext() context.setParameter(VadParameter.Sensitivity, 0.8) @@ -40,7 +38,7 @@ async function main() { console.log('Sensitivity:', context.getParameter(VadParameter.Sensitivity)) console.log('Prediction delay:', context.getPredictionDelay(), 'samples') - // Silence, so nothing should be reported. Feed real speech to see this flip. + // Process silence as sample input. Replace this with audio from your source. let audio = new Float32Array(blockSize) for (let block = 0; block < 10; block += 1) { // Resolves to the same samples, unmodified, so one variable carries the stream. @@ -50,26 +48,19 @@ async function main() { console.log('\nSpeech detected:', context.isSpeechDetected()) console.log('Raw probability:', context.getRawVadProbability().toFixed(4)) - // Detection and enhancement together. - // - // The VAD must see the *original* audio, not the processor's output: enhancement is - // designed to change the signal, so detecting on its output runs the VAD model on audio - // it was not trained for, and stacks the processor's delay onto the prediction. The VAD - // hands its block back untouched, so ordering the two calls is enough. + // Run detection before enhancement so the VAD receives the original input. + // Enhanced audio changes the signal seen by the VAD and adds processing delay. console.log('\nRunning detection and enhancement on the same stream') const enhancementModel = Model.fromFile(await Model.download(ENHANCEMENT_MODEL_ID, MODEL_DIR)) - // Initialized with the *VAD's* block size, not the enhancement model's own optimum, so - // that one block can be handed to both. Two models need not agree on an optimal block - // size, and feeding a processor a block it was not configured for is an error, so when - // sharing a stream, pick one size and configure everything with it. + // Configure both instances with the same block size to share input blocks. + // Models may have different optimal block sizes. const processor = await new ProcessorAsync(enhancementModel, licenseKey).withConfig(sampleRate, blockSize) const processorContext = await processor.getContext() for (let i = 0; i < 10; i += 1) { - // In a real stream this block comes from the source. What matters is that it is never - // the previous iteration's output: enhanced audio must not come back round into the VAD. + // Read a new input block for each iteration; do not reuse enhanced output as VAD input. const block = Float32Array.from({ length: blockSize }, () => (Math.random() - 0.5) * 0.2) // The VAD sees the input, then the processor enhances it. @@ -79,8 +70,7 @@ async function main() { console.log(`Block ${i}: speech ${context.isSpeechDetected()}, ${enhanced.length} samples out`) } - // The two delays describe different things and are independent: one shifts the audio, - // the other tells you how far behind the decision is. + // Audio delay and VAD prediction delay are independent measurements in samples. console.log('Audio delay:', processorContext.getAudioDelay(), 'samples') console.log('Prediction delay:', context.getPredictionDelay(), 'samples') diff --git a/examples/vad.js b/examples/vad.js index f8f73c5..8169351 100644 --- a/examples/vad.js +++ b/examples/vad.js @@ -45,10 +45,10 @@ async function main() { console.log('Speech hold duration:', context.getParameter(VadParameter.SpeechHoldDuration), 's') // The decision lags its input by this many samples. It is not applied to the audio, so - // use it to line speech decisions up with the audio timeline. + // use it to align speech decisions with the input audio. console.log('Prediction delay:', context.getPredictionDelay(), 'samples') - // Silence, so nothing should be reported. Feed real speech to see this flip. + // Process silence as sample input. Replace this with audio from your source. const audio = new Float32Array(blockSize) for (let block = 0; block < 10; block += 1) { // Reads the block without modifying it, so the same buffer can go on to a processor. diff --git a/index.d.ts b/index.d.ts index 425c646..305c824 100644 --- a/index.d.ts +++ b/index.d.ts @@ -1,281 +1,333 @@ /* auto-generated by NAPI-RS */ /* eslint-disable */ /** - * Analyzer for analysis models such as Tyto. + * Analyzes audio quality using an analysis model, such as Tyto. * - * Buffering and analysis are separate calls: {@link Analyzer#buffer} is cheap enough for - * the audio path, while running the model is not. Analysis comes in two forms, - * {@link Analyzer#analyzeAsync} on a worker thread and {@link Analyzer#analyze} on the - * calling thread. + * Call {@link Analyzer#initialize}, then {@link Analyzer#buffer} to collect mono audio. + * The model determines how much audio is retained; older samples are discarded as new + * audio arrives. * - * Only a fixed span of audio is retained, determined by the model; older audio is - * discarded as more is buffered. - * - * The SDK splits this into a collector and an analyzer so the two halves can live on - * different threads. A class instance cannot cross into a Node worker, so both are - * exposed as one object here, and the split shows through: {@link Analyzer#buffer} drives - * the collector on the calling thread, while {@link Analyzer#analyzeAsync} moves the - * analyzer half onto a worker. The SDK guarantees the two are safe to use concurrently. + * Run {@link Analyzer#analyzeAsync} to analyze the buffered audio on a libuv worker + * thread, or {@link Analyzer#analyze} to run it on the calling thread. Analysis is + * computationally expensive and should not run in audio processing callbacks. + * Buffering remains synchronous and can continue while async analysis is running. */ export declare class Analyzer { - /** Creates an analyzer from an analysis model. Other model types are rejected. */ + /** + * Creates an analyzer from an analysis model. + * + * Construction allocates memory. Call {@link Analyzer#initialize} before buffering audio. + * + * @param model - Analysis model. Other model types are rejected. + * @param licenseKey - SDK license key from . + */ constructor(model: Model, licenseKey: string) /** - * Destroys the native collector and analyzer immediately, releasing their memory - * without waiting for garbage collection. + * Destroys the native collector and analyzer without waiting for garbage collection. * - * Every later method throws; calling `dispose()` again does nothing. Blocks until an - * in-flight `analyzeAsync` on a worker thread finishes. + * After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. + * This call blocks the calling thread while a worker holds the analyzer lock. + * Queued analysis that acquires the lock after disposal rejects its promise. */ dispose(): void /** - * Configures the analyzer for an audio format. Must be called before buffering. + * Configures the analyzer's audio collector. + * + * Call this before buffering audio. Use {@link Model#getOptimalSampleRate} and + * {@link Model#getOptimalBlockSize} to avoid internal resampling and rebuffering. + * This method allocates memory; avoid calling it from audio processing callbacks. * - * The model's optimal sample rate and block size avoid internal resampling and - * rebuffering. Allocates, so keep it off the audio path. + * @param sampleRate - Audio sample rate in Hz. + * @param blockSize - Number of mono samples per block. + * @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + * Blocks larger than `blockSize` are always rejected. */ initialize(sampleRate: number, blockSize: number, variableBlockSize?: boolean | undefined | null): void - /** Buffers a mono audio block for later analysis, leaving the audio unmodified. */ - buffer(audio: Float32Array): void /** - * Runs the analysis model over the buffered audio, on the calling thread. + * Buffers a mono audio block for later analysis without modifying the input. * - * The model consumes a fixed span of audio. Calling this before that much has been - * buffered analyzes what is there, padded with silence. + * Call {@link Analyzer#initialize} first. The block must contain exactly `blockSize` + * samples, or at most `blockSize` if `variableBlockSize` is enabled. + * Buffering does not acquire the analyzer lock and can run during async analysis. + */ + buffer(audio: Float32Array): void + /** + * Analyzes buffered audio on the calling thread and returns the scores. * - * Analysis is mono. Mix multichannel audio down, or use one analyzer per channel. + * The model analyzes a fixed duration of audio. If less audio has been buffered, the + * remaining input is padded with silence. * - * This call is expensive and blocks. Prefer {@link Analyzer#analyzeAsync} unless - * nothing else is waiting on the event loop. + * This method blocks the calling thread. Use {@link Analyzer#analyzeAsync} to keep the + * event loop available. For multichannel audio, mix down to mono or use one analyzer + * per channel. */ analyze(): AnalysisResult /** - * Runs the analysis model over the buffered audio on a worker thread. + * Analyzes buffered audio on a libuv worker thread and returns a promise for the scores. * - * Same result as {@link Analyzer#analyze}, off the event loop. Prefer this wherever - * the analysis model is too expensive to run on the calling thread. + * Uses the same analysis and silence padding as {@link Analyzer#analyze}. + * {@link Analyzer#buffer} can continue collecting audio while analysis runs. * - * {@link Analyzer#buffer} stays synchronous and takes no lock, so audio can keep - * arriving while an analysis is in flight; the SDK guarantees the collector and - * analyzer halves are safe to use concurrently. The other methods here do take the - * analyzer's lock, so calling {@link Analyzer#analyze}, {@link Analyzer#reset} or - * {@link Analyzer#terminateSession} while this is pending blocks the calling thread - * until it finishes. + * Synchronous calls to `analyze`, `reset`, `updateBearerToken`, `terminateSession` and + * `dispose` wait for the analyzer lock and may block while analysis is running. */ analyzeAsync(): Promise - /** Clears buffered audio and internal state, keeping the configured audio settings. */ + /** Clears buffered audio and internal state while preserving the configured audio settings. */ reset(): void /** - * Swaps in a renewed JWT without tearing down the analyzer. + * Replaces the bearer token on the running analyzer. + * + * Use this to refresh a JWT without recreating the instance. Both the original license + * key and the new token must be JWTs. If this call fails, the previous token remains active. + * + * A successful call validates the token's format and applies it immediately. Backend + * acceptance is checked later. If the backend rejects the token, the SDK retries with + * backoff; analysis may be rejected if no accepted token arrives in time. + * Supply a valid token to recover the session. * - * Only works when both the original key and the new token are JWTs. On failure the - * call is a no-op and the previous token stays active. + * This method allocates memory and takes a mutex. Avoid calling it from audio processing callbacks. */ updateBearerToken(token: string): void - /** Ends this analyzer's telemetry session, after which it can no longer analyze audio. */ + /** + * Terminates the telemetry session associated with this analyzer. + * + * Once termination is handled, the analyzer can no longer analyze buffered audio. + * The session also ends when the native analyzer is destroyed. This method may block; + * avoid calling it from audio processing callbacks. If another session is still active, + * termination can complete asynchronously. + */ terminateSession(): void } /** * A loaded ai-coustics model. * - * One model can back multiple processors, VADs and analyzers, according to its type. - * The underlying model data is kept alive by every object created from it through - * internal reference counting, so this handle may be released first. + * The same model can be used to create multiple independent instances of a compatible + * processor, VAD or analyzer. Each instance retains a reference to the model data, + * so the `Model` handle can be disposed or garbage collected first. */ export declare class Model { /** * Loads a model from a `.aicmodel` file. * - * The model data is memory-mapped, not copied into the process, so the file must - * not be modified or deleted while this model, or any object created from it, - * is alive. + * The SDK memory-maps the file. Do not modify or delete it while the model or any + * processor, VAD or analyzer created from it is still alive. * - * Browse available models at , or fetch one with - * {@link Model.download}. + * Download models with {@link Model.download}. Available model IDs are listed at + * . + * + * @param path - Path to the model file. + * @throws If the file cannot be loaded or its format is incompatible with this SDK. */ static fromFile(path: string): Model /** - * Unmaps the model file immediately, releasing its footprint without waiting for - * garbage collection. + * Releases this handle's reference to the native model without waiting for garbage collection. + * + * Instances created from this model remain usable. The file mapping is released when + * its last reference is destroyed. * - * Objects already created from the model keep working: the SDK keeps the weights - * alive through internal reference counting. Every later method on this handle - * throws; calling `dispose()` again does nothing. + * After disposal, all methods except `dispose()` throw. Repeated disposal has no effect. */ dispose(): void /** - * Downloads a model from the ai-coustics artifact CDN and resolves to its path. + * Downloads a model from the ai-coustics CDN and returns a promise for its file path. * - * The manifest is re-fetched on every call so the newest compatible model version - * is always used. An existing file with a matching checksum is not re-downloaded; - * one with a mismatching checksum is replaced. + * Each call fetches the manifest to select the latest compatible model version. + * An existing file is reused if its checksum matches; otherwise it is replaced. + * The download runs on Node's libuv thread pool. * - * The download runs on a worker thread, so it does not block the event loop. + * @param modelId - Model ID listed at . + * @param downloadDir - Directory in which to store the model. + * @returns A promise for the downloaded or cached model's path. */ static download(modelId: string, downloadDir: string): Promise - /** The model identifier, e.g. `quail-vf-2.2-s-16khz`. */ + /** Returns the model identifier, including its build and format version suffixes. */ getId(): string /** - * The sample rate in Hz the model was trained for. + * Returns the sample rate in Hz for which the model was trained. * - * Audio at any rate can be processed, but a model only enhances frequencies up to - * its own Nyquist limit, so matching this rate gives the best quality. + * The SDK resamples audio at other supported rates internally. Enhancement is limited + * to frequencies below half the model's sample rate. */ getOptimalSampleRate(): number /** - * The block size that avoids internal buffering at `sampleRate`. + * Returns the optimal block size in samples for the given sample rate. * - * Any other block size adds buffering latency on top of the base processing delay. - * The value changes with the sample rate, because the model works on a fixed time - * window: a 10 ms window is 480 samples at 48 kHz but 160 at 16 kHz. + * Use this value to avoid additional buffering latency. The block size depends on the + * sample rate because each model processes a fixed duration of audio. + * + * @param sampleRate - Audio sample rate in Hz. */ getOptimalBlockSize(sampleRate: number): number } /** - * Speech enhancement processor. + * Processes mono audio with an enhancement or bypass model. * - * Built from an enhancement or bypass model. Use {@link Vad} for dedicated VAD models - * and {@link Analyzer} for analysis models; passing the wrong kind throws. + * Call {@link Processor#initialize} before processing audio. Each processor maintains + * its own state; create one instance per audio stream. Multiple processors can share + * the same {@link Model}. * - * Create several processors to handle multiple streams or to switch models at runtime. + * Use {@link ProcessorAsync} to run processing on Node's libuv thread pool. */ export declare class Processor { /** - * Creates a processor from an enhancement or bypass model. + * Creates a new speech enhancement processor. + * + * Construction is synchronous and throws if creation fails. Call + * {@link Processor#initialize} before processing audio. * - * Telemetry follows the runtime environment; pass `otelConfig` to override it for this - * instance. + * @param model - Enhancement or bypass model. Other model types are rejected. + * @param licenseKey - SDK license key from . + * @param otelConfig - Optional telemetry configuration. When omitted, telemetry follows + * the runtime environment. */ constructor(model: Model, licenseKey: string, otelConfig?: OtelConfig | undefined | null) /** - * Destroys the native processor immediately, releasing its memory and telemetry - * session without waiting for garbage collection. + * Destroys the native processor and releases its telemetry session. * - * Every later method throws; calling `dispose()` again does nothing. + * Use this for cleanup at a specific point instead of waiting for garbage collection. + * After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. */ dispose(): void /** - * Configures the processor for an audio format. Must be called before processing. + * Configures the processor for the given audio format. * - * For the lowest delay use {@link Model#getOptimalSampleRate} and - * {@link Model#getOptimalBlockSize}. Allocates, so keep it off the audio path. + * Call this method before processing audio. Use {@link Model#getOptimalSampleRate} and + * {@link Model#getOptimalBlockSize} for the lowest delay. + * This method allocates memory; avoid calling it from audio processing callbacks. * - * With `variableBlockSize` enabled (default `false`), calls shorter than `blockSize` - * are permitted at the cost of extra delay; longer calls are always rejected. + * @param sampleRate - Audio sample rate in Hz. + * @param blockSize - Number of mono samples per block. + * @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + * Variable block sizes can add buffering latency. Larger blocks are always rejected. */ initialize(sampleRate: number, blockSize: number, variableBlockSize?: boolean | undefined | null): void /** * Enhances a mono audio block in place. * - * The block must be exactly `blockSize` samples, or at most `blockSize` if - * `variableBlockSize` was enabled. + * Call {@link Processor#initialize} first. The block must contain exactly `blockSize` + * samples, or at most `blockSize` if `variableBlockSize` is enabled. + * If the input uses a SharedArrayBuffer, prevent other workers from accessing it during + * this call. */ process(audio: Float32Array): void /** * Creates a handle for reading and writing this processor's parameters and state. * - * Each call returns an independent handle onto the same processor. + * Each call returns an independent handle to the same processor. */ getContext(): ProcessorContext /** - * Ends this processor's telemetry session, after which it can no longer process audio. + * Terminates the telemetry session associated with this processor. * - * A session is closed automatically when the processor is collected, but GC timing is - * not guaranteed, so call this on a lifecycle event instead. May block, so keep it off - * the audio path. + * Once termination is handled, the processor can no longer process audio. + * The session also ends when the native object is destroyed. Use this method when + * termination must be requested at a specific lifecycle event. + * + * This method may block. Avoid calling it from audio processing callbacks. + * If another session is still active, termination can complete asynchronously. */ terminateSession(): void } /** - * Speech enhancement processor that keeps its work off the main thread. + * Speech enhancement for use in async applications. * - * The same processing as {@link Processor}, but each call returns a promise and runs on - * Node's libuv thread pool, so the event loop stays responsive. Prefer this when other - * work shares that loop, as in a server; prefer {@link Processor} on a dedicated audio - * thread or in a batch script, where nothing else needs the loop. + * Initialization, processing and context creation run on Node's libuv thread pool and + * return promises. Construction and disposal are synchronous. * - * Mirrors `ProcessorAsync` in the Rust SDK. + * Use {@link Processor} when processing should run on the calling thread. * - * ### Concurrency + * ### Threading * - * One instance handles one stream. Do not start a second {@link ProcessorAsync#process} - * before the first resolves: libuv completes work items out of order, which would - * desync the stream. To process several streams at once, create several instances. + * Use one instance per stream and await each operation before submitting the next. + * Concurrent calls on one instance are not guaranteed to execute in submission order. + * Use separate instances to process multiple streams concurrently. * - * The libuv pool is four threads by default and is shared with `fs`, `dns` and `crypto`. - * Raise `UV_THREADPOOL_SIZE` before Node starts to run more streams in parallel. - * `AIC_NUM_THREADS` has no effect: it sizes a rayon pool this binding does not use. + * The libuv pool defaults to four threads and is shared with filesystem, DNS and crypto + * work. Set `UV_THREADPOOL_SIZE` before starting Node to change its size. + * `AIC_NUM_THREADS` does not apply to these bindings. */ export declare class ProcessorAsync { /** - * Creates a processor from an enhancement or bypass model. + * Creates a new async speech enhancement processor. * - * Construction is synchronous and throws on failure, as in the Rust SDK; only the - * audio work runs on a worker thread. + * Construction is synchronous and throws if creation fails. Await + * {@link ProcessorAsync#initialize} or {@link ProcessorAsync#withConfig} before processing audio. * - * Telemetry follows the runtime environment; pass `otelConfig` to override it for this - * instance. + * @param model - Enhancement or bypass model. Other model types are rejected. + * @param licenseKey - SDK license key from . + * @param otelConfig - Optional telemetry configuration. When omitted, telemetry follows + * the runtime environment. */ constructor(model: Model, licenseKey: string, otelConfig?: OtelConfig | undefined | null) /** - * Destroys the native processor immediately, releasing its memory and telemetry - * session without waiting for garbage collection. + * Destroys the native processor and releases its telemetry session. + * + * Use this for cleanup at a specific point instead of waiting for garbage collection. + * After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. * - * Every later method throws; calling `dispose()` again does nothing. Blocks until - * in-flight work on the libuv pool finishes. + * This call blocks the calling thread while a worker holds the instance lock. + * Queued work that acquires the lock after disposal rejects its promise. */ dispose(): void /** - * Initializes the processor and resolves to a handle onto it, for chaining off the - * constructor: + * Initializes the processor and returns a promise for a handle to the initialized instance. * - * ```js - * const processor = await new ProcessorAsync(model, licenseKey).withConfig(48000, 480) - * ``` + * Uses the same configuration as {@link ProcessorAsync#initialize}. The returned handle and + * this object share the same native instance; disposing either invalidates both. * - * The handle it resolves to drives the same underlying processor as the receiver, so - * either one can be used afterwards. The Rust SDK returns `self` here, which a promise - * cannot express. + * ```javascript + * const processor = await new ProcessorAsync(model, licenseKey).withConfig(sampleRate, blockSize) + * ``` */ withConfig(sampleRate: number, blockSize: number, variableBlockSize?: boolean | undefined | null): Promise /** - * Configures the processor for an audio format. Must be called before processing. + * Configures the processor for the given audio format. + * + * Await this method before processing audio. Use {@link Model#getOptimalSampleRate} and + * {@link Model#getOptimalBlockSize} for the lowest delay. + * Initialization allocates memory and runs on a libuv worker thread. * - * See {@link Processor#initialize}. Allocates, so it runs on a worker. + * @param sampleRate - Audio sample rate in Hz. + * @param blockSize - Number of mono samples per block. + * @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + * Variable block sizes can add buffering latency. Larger blocks are always rejected. */ initialize(sampleRate: number, blockSize: number, variableBlockSize?: boolean | undefined | null): Promise /** - * Enhances a mono audio block and resolves to the enhanced samples. + * Enhances a mono audio block and returns a promise for the enhanced samples. * - * Unlike {@link Processor#process} this does **not** write into the caller's array. - * The samples are copied out before the work is queued, so the input stays valid and - * untouched while the promise is pending, and the result arrives as a new array: + * The input is copied before work is queued and remains unmodified. The promise + * resolves to a new `Float32Array` containing the enhanced samples. * - * ```js - * let audio = new Float32Array(blockSize) - * for (;;) audio = await processor.process(audio) - * ``` + * The instance must be initialized first. The block must contain exactly `blockSize` + * samples, or at most `blockSize` if `variableBlockSize` is enabled. + * Await each call before submitting the next block. * - * The block must be exactly `blockSize` samples, or at most `blockSize` if - * `variableBlockSize` was enabled. + * ```javascript + * const enhanced = await processor.process(block) + * ``` */ process(audio: Float32Array): Promise> /** - * Creates a handle for reading and writing this processor's parameters and state. + * Returns a promise for a {@link ProcessorContext} to control this processor. * - * Asynchronous because it takes the processor lock, which a queued `process` may - * briefly hold; awaiting keeps that wait off the event loop. The returned handle is - * the same {@link ProcessorContext} the synchronous class hands out, with the same - * synchronous methods. + * Context creation runs on a worker thread because it may wait for processing to + * release the instance lock. The returned context's methods are synchronous and can + * be called while audio is being processed. */ getContext(): Promise /** - * Ends this processor's telemetry session, after which it can no longer process audio. + * Terminates the telemetry session associated with this processor. + * + * Once termination is handled, the processor can no longer process audio. + * The session also ends when the native object is destroyed. Use this method when + * termination must be requested at a specific lifecycle event. * - * May block, so it runs on a worker. + * Termination runs on a libuv worker thread because it may block. + * If another session is still active, termination can complete asynchronously. */ terminateSession(): Promise } @@ -291,170 +343,206 @@ export declare class ProcessorAsync { export declare class ProcessorContext { /** Sets an enhancement parameter. Throws if the value is out of range. */ setParameter(parameter: ProcessorParameter, value: number): void - /** Reads the current value of a parameter. */ + /** Returns the current value of an enhancement parameter. */ getParameter(parameter: ProcessorParameter): number /** - * Total delay the processor applies to the audio, in samples at the initialized rate. + * Returns the audio delay in samples at the configured sample rate. * * Covers algorithmic delay plus any buffering from a non-optimal block size. Before * initialization it reports the base delay at the model's optimal settings. */ getAudioDelay(): number /** - * Clears internal state and buffers, keeping the configured audio settings. + * Clears internal state and buffers while preserving the configured audio settings. * - * Call this on a stream discontinuity or when seeking, to keep earlier audio from - * bleeding into the output. + * Call this when the stream is interrupted or when seeking to prevent previous audio + * from affecting the output. */ reset(): void /** - * Swaps in a renewed JWT without interrupting processing. + * Replaces the bearer token on the running processor. + * + * Use this to refresh a JWT without recreating the instance. Both the original license + * key and the new token must be JWTs. If this call fails, the previous token remains active. * - * Only works when both the original key and the new token are JWTs. On failure the - * call is a no-op and the previous token stays active. On success the swap is applied - * immediately and is not gated on backend acceptance. + * A successful call validates the token's format and applies it immediately. Backend + * acceptance is checked later. If the backend rejects the token, the SDK retries with + * backoff; processing is eventually disabled if no accepted token arrives in time. + * Supply a valid token to recover the session. + * + * This method allocates memory and takes a mutex. Avoid calling it from audio processing callbacks. */ updateBearerToken(token: string): void } /** - * Voice activity detector running a dedicated VAD model. + * Detects speech using a dedicated VAD model. * - * Driven explicitly through {@link Vad#process} and independent of any - * {@link Processor}; predictions are read through a {@link VadContext}. Enhancement - * models are rejected. + * Call {@link Vad#initialize}, then pass mono audio to {@link Vad#process}. + * Processing leaves the audio unmodified and updates the prediction, which can be read + * through a {@link VadContext}. * - * When enhancement and detection run together, feed this the **original** audio, not the - * processor's output: enhancement changes the signal the VAD model expects, and stacks - * the processor's delay onto the prediction. `process` leaves its input untouched, so - * call it on the same block before `Processor#process`. + * When using enhancement and detection together, pass the original input to the VAD + * before calling {@link Processor#process}. Enhanced audio changes the signal seen by + * the VAD and adds the processor's audio delay to the prediction delay. */ export declare class Vad { /** - * Creates a voice activity detector from a dedicated VAD model. + * Creates a new voice activity detector. + * + * Construction is synchronous and throws if creation fails. Call + * {@link Vad#initialize} before processing audio. * - * Telemetry follows the runtime environment; pass `otelConfig` to override it for this - * instance. + * @param model - Dedicated VAD model. Other model types are rejected. + * @param licenseKey - SDK license key from . + * @param otelConfig - Optional telemetry configuration. When omitted, telemetry follows + * the runtime environment. */ constructor(model: Model, licenseKey: string, otelConfig?: OtelConfig | undefined | null) /** - * Destroys the native VAD immediately, releasing its memory and telemetry session - * without waiting for garbage collection. + * Destroys the native VAD and releases its telemetry session. * - * Every later method throws; calling `dispose()` again does nothing. + * Use this for cleanup at a specific point instead of waiting for garbage collection. + * After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. */ dispose(): void /** - * Configures the VAD for an audio format. Must be called before processing. + * Configures the VAD for the given audio format. * - * The model's optimal sample rate and block size give the most frequent prediction - * updates. Allocates, so keep it off the audio path. + * Call this method before processing audio. Use {@link Model#getOptimalSampleRate} and + * {@link Model#getOptimalBlockSize} for the most frequent prediction updates. + * This method allocates memory; avoid calling it from audio processing callbacks. + * + * @param sampleRate - Audio sample rate in Hz. + * @param blockSize - Number of mono samples per block. + * @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + * Variable block sizes can add buffering latency. Larger blocks are always rejected. */ initialize(sampleRate: number, blockSize: number, variableBlockSize?: boolean | undefined | null): void - /** Examines a mono audio block and updates the prediction, leaving the audio unmodified. */ + /** + * Processes a mono audio block and updates the VAD prediction without modifying the input. + * + * Call {@link Vad#initialize} first. The block must contain exactly `blockSize` samples, + * or at most `blockSize` if `variableBlockSize` is enabled. + */ process(audio: Float32Array): void /** * Creates a handle for reading predictions and controlling this VAD. * - * Each call returns an independent handle onto the same VAD. + * Each call returns an independent handle to the same VAD. */ getContext(): VadContext - /** Ends this VAD's telemetry session, after which it can no longer process audio. */ + /** + * Terminates the telemetry session associated with this VAD. + * + * Once termination is handled, the VAD can no longer process audio. + * The session also ends when the native object is destroyed. Use this method when + * termination must be requested at a specific lifecycle event. + * + * This method may block. Avoid calling it from audio processing callbacks. + * If another session is still active, termination can complete asynchronously. + */ terminateSession(): void } /** - * Voice activity detector that keeps its work off the main thread. + * Voice activity detection for use in async applications. * - * The same detection as {@link Vad}, but each call returns a promise and runs on Node's - * libuv thread pool, so the event loop stays responsive. Predictions are read through a - * {@link VadContext}, whose methods are all synchronous. + * Initialization, processing and context creation run on Node's libuv thread pool and + * return promises. Construction and disposal are synchronous. * - * Mirrors `VadAsync` in the Rust SDK. + * Read predictions through a {@link VadContext}. Pass the original input audio to the + * VAD before enhancement. * - * As with {@link Vad}, feed this the **original** audio when enhancement and detection - * run together, not a processor's output. + * ### Threading * - * ### Concurrency + * Use one instance per stream and await each operation before submitting the next. + * Concurrent calls on one instance are not guaranteed to execute in submission order. + * Use separate instances to process multiple streams concurrently. * - * One instance handles one stream. Do not start a second {@link VadAsync#process} before - * the first resolves: libuv completes work items out of order, which would desync the - * stream and scramble the prediction. To watch several streams at once, create several - * instances. - * - * The libuv pool is four threads by default and is shared with `fs`, `dns` and `crypto`. - * Raise `UV_THREADPOOL_SIZE` before Node starts to run more streams in parallel. - * `AIC_NUM_THREADS` has no effect: it sizes a rayon pool this binding does not use. + * The libuv pool defaults to four threads and is shared with filesystem, DNS and crypto + * work. Set `UV_THREADPOOL_SIZE` before starting Node to change its size. + * `AIC_NUM_THREADS` does not apply to these bindings. */ export declare class VadAsync { /** - * Creates a voice activity detector from a dedicated VAD model. + * Creates a new async voice activity detector. * - * Construction is synchronous and throws on failure, as in the Rust SDK; only the - * audio work runs on a worker thread. + * Construction is synchronous and throws if creation fails. Await + * {@link VadAsync#initialize} or {@link VadAsync#withConfig} before processing audio. * - * Telemetry follows the runtime environment; pass `otelConfig` to override it for this - * instance. + * @param model - Dedicated VAD model. Other model types are rejected. + * @param licenseKey - SDK license key from . + * @param otelConfig - Optional telemetry configuration. When omitted, telemetry follows + * the runtime environment. */ constructor(model: Model, licenseKey: string, otelConfig?: OtelConfig | undefined | null) /** - * Destroys the native VAD immediately, releasing its memory and telemetry session - * without waiting for garbage collection. + * Destroys the native VAD and releases its telemetry session. + * + * Use this for cleanup at a specific point instead of waiting for garbage collection. + * After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. * - * Every later method throws; calling `dispose()` again does nothing. Blocks until - * in-flight work on the libuv pool finishes. + * This call blocks the calling thread while a worker holds the instance lock. + * Queued work that acquires the lock after disposal rejects its promise. */ dispose(): void /** - * Initializes the VAD and resolves to a handle onto it, for chaining off the - * constructor: + * Initializes the VAD and returns a promise for a handle to the initialized instance. * - * ```js - * const vad = await new VadAsync(model, licenseKey).withConfig(16000, 160) - * ``` + * Uses the same configuration as {@link VadAsync#initialize}. The returned handle and + * this object share the same native instance; disposing either invalidates both. * - * The handle it resolves to drives the same underlying VAD as the receiver, so either - * one can be used afterwards. The Rust SDK returns `self` here, which a promise cannot - * express. + * ```javascript + * const vad = await new VadAsync(model, licenseKey).withConfig(sampleRate, blockSize) + * ``` */ withConfig(sampleRate: number, blockSize: number, variableBlockSize?: boolean | undefined | null): Promise /** - * Configures the VAD for an audio format. Must be called before processing. + * Configures the VAD for the given audio format. + * + * Await this method before processing audio. Use {@link Model#getOptimalSampleRate} and + * {@link Model#getOptimalBlockSize} for the most frequent prediction updates. + * Initialization allocates memory and runs on a libuv worker thread. * - * See {@link Vad#initialize}. Allocates, so it runs on a worker. + * @param sampleRate - Audio sample rate in Hz. + * @param blockSize - Number of mono samples per block. + * @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + * Variable block sizes can add buffering latency. Larger blocks are always rejected. */ initialize(sampleRate: number, blockSize: number, variableBlockSize?: boolean | undefined | null): Promise /** - * Examines a mono audio block, updates the prediction, and resolves to the same - * samples unmodified. + * Updates the VAD prediction and returns a promise for the original mono audio samples. * - * The samples are copied out before the work is queued, so the caller's array stays - * valid and untouched while the promise is pending. The block is handed back, instead - * of the promise resolving to nothing, to match the Rust SDK and to keep a streaming - * loop reading the same either side of the boundary: + * The input is copied before work is queued and remains unmodified. The promise + * resolves to a new `Float32Array` containing the original samples. * - * ```js - * let audio = new Float32Array(blockSize) - * for (;;) { - * audio = await vad.process(audio) - * console.log(context.isSpeechDetected()) - * } + * The instance must be initialized first. The block must contain exactly `blockSize` + * samples, or at most `blockSize` if `variableBlockSize` is enabled. + * Await each call before submitting the next block. + * + * ```javascript + * const audio = await vad.process(block) * ``` */ process(audio: Float32Array): Promise> /** - * Creates a handle for reading predictions and controlling this VAD. + * Returns a promise for a {@link VadContext} to control this VAD and read predictions. * - * Asynchronous because it takes the VAD lock, which a queued `process` may briefly - * hold; awaiting keeps that wait off the event loop. The returned handle is the same - * {@link VadContext} the synchronous class hands out, whose methods are synchronous, - * so a prediction can be read from inside an audio callback. + * Context creation runs on a worker thread because it may wait for processing to + * release the instance lock. The returned context's methods are synchronous and can + * be called while audio is being processed. */ getContext(): Promise /** - * Ends this VAD's telemetry session, after which it can no longer process audio. + * Terminates the telemetry session associated with this VAD. * - * May block, so it runs on a worker. + * Once termination is handled, the VAD can no longer process audio. + * The session also ends when the native object is destroyed. Use this method when + * termination must be requested at a specific lifecycle event. + * + * Termination runs on a libuv worker thread because it may block. + * If another session is still active, termination can complete asynchronously. */ terminateSession(): Promise } @@ -468,158 +556,172 @@ export declare class VadAsync { export declare class VadContext { /** Sets a VAD parameter. Throws if the value is out of range. */ setParameter(parameter: VadParameter, value: number): void - /** Reads the current value of a VAD parameter. */ + /** Returns the current value of a VAD parameter. */ getParameter(parameter: VadParameter): number /** - * Whether speech is currently detected. + * Returns whether speech is currently detected. * * The decision lags its input by {@link VadContext#getPredictionDelay} samples, and * stops updating if the backing VAD stops being processed. */ isSpeechDetected(): boolean /** - * The model's raw prediction, in the range 0.0 - 1.0. + * Returns the model's speech probability in the range 0.0 to 1.0. * - * Unlike {@link VadContext#isSpeechDetected} this skips the SDK's post-processing - * (speech hold, sensitivity thresholding), for building your own abstractions on top. - * The same latency notes apply. + * This value excludes speech hold, sensitivity thresholding and minimum speech duration. + * Use it to implement custom detection logic. The prediction delay reported by + * {@link VadContext#getPredictionDelay} also applies to this value. */ getRawVadProbability(): number /** - * How far the prediction lags its input, in samples at the initialized rate. + * Returns the prediction delay in samples at the configured sample rate. + * + * Includes input buffering, STFT and model processing. Non-optimal or variable block + * sizes can add buffering latency. Convert to milliseconds with + * `delaySamples * 1000 / sampleRate`. * - * Covers input reblocking, STFT and model processing. This delay is **not** applied to - * the audio (`process` leaves the buffer untouched), so use it to line speech - * decisions up with the audio timeline. Independent of a processor's audio delay. + * Use this delay to align speech decisions with the input audio. It is independent of + * a processor's audio delay; VAD processing does not delay or modify the audio. */ getPredictionDelay(): number /** - * Clears internal state, including the published prediction. + * Clears internal state and buffers, including the published speech decision and probability. * - * Call this on a stream discontinuity or when seeking, to keep earlier audio from - * causing mispredictions. + * Call this when the stream is interrupted or when seeking to prevent predictions + * from using previous audio. The VAD remains initialized with its configured settings. */ reset(): void /** - * Swaps in a renewed JWT without interrupting processing. + * Replaces the bearer token on the running VAD. * - * Only works when both the original key and the new token are JWTs. On failure the - * call is a no-op and the previous token stays active. + * Use this to refresh a JWT without recreating the instance. Both the original license + * key and the new token must be JWTs. If this call fails, the previous token remains active. + * + * A successful call validates the token's format and applies it immediately. Backend + * acceptance is checked later. If the backend rejects the token, the SDK retries with + * backoff; processing is eventually disabled if no accepted token arrives in time. + * Supply a valid token to recover the session. + * + * This method allocates memory and takes a mutex. Avoid calling it from audio processing callbacks. */ updateBearerToken(token: string): void } /** - * Overrides the telemetry wrapper id. Internal only, for ai-coustics wrappers embedding + * Overrides the telemetry wrapper ID. Internal only, for ai-coustics wrappers embedding * this package (e.g. the LiveKit plugin): call before constructing any `Processor`, `Vad` - * or `Analyzer`, whose constructors otherwise claim the id for this SDK. The id can only + * or `Analyzer`, whose constructors otherwise claim the ID for this SDK. The ID can only * be set once per process; later writes are silently discarded. */ export declare function _setSdkId(id: number): void /** - * Scores produced by {@link Analyzer#analyze}. + * Results of analyzing an audio signal with {@link Analyzer}. * - * Every score runs 0.0 - 1.0. For all of them except `speakerLoudness`, lower means less - * problematic audio. + * Scores range from 0.0 to 1.0. For every field except `speakerLoudness`, lower values + * indicate less problematic audio. */ export interface AnalysisResult { /** - * Headline score: how likely this audio is to break downstream models such as - * speech-to-text, VAD, turn-taking or speech-to-speech. + * Predicts the likelihood of failure in downstream models, including speech-to-text, + * voice activity detection, turn-taking and speech-to-speech models. */ riskScore: number - /** How distant and reverberant the speaker sounds. */ + /** Measure of speaker distance and reverberation. */ speakerReverb: number - /** How loud the speaker is. */ + /** Measure of speaker loudness. */ speakerLoudness: number - /** How much speech from people other than the main speaker is present. */ + /** Measure of interfering speech from sources other than the main speaker. */ interferingSpeech: number - /** How much ambient or environmental noise is present. */ + /** Measure of ambient or environmental noise. */ noise: number - /** Artifacts from lossy speech codecs, e.g. a low bitrate or narrowband codec. */ + /** Measure of artifacts from lossy speech codecs, such as low bitrate or narrowband codecs. */ codecDegradation: number /** - * Dropouts and discontinuities, e.g. from packet loss, frame erasure, jitter or CPU + * Measure of audio dropouts and discontinuities from packet loss, frame erasure, jitter or CPU * overload. */ packetLoss: number } -/** - * The model file format version this SDK can load. - * - * Model URLs are versioned by this number, so it decides which model files are usable. - */ +/** Returns the model file format version supported by this SDK. */ export declare function getCompatibleModelVersion(): number /** - * The version of the underlying native SDK, e.g. `"0.23.0"`. + * Returns the version of the underlying native SDK. * - * Not necessarily this package's version. + * This may differ from the Node.js package version. */ export declare function getVersion(): string /** - * Per-instance OpenTelemetry settings. + * OpenTelemetry configuration for a processor or VAD instance. * - * Overrides the environment-based defaults (e.g. `AIC_SDK_OTEL_ENABLE`) for the one - * processor or VAD it is passed to. + * Overrides the SDK's environment-based telemetry settings, such as + * `AIC_SDK_OTEL_ENABLE`, for the instance receiving this configuration. */ export interface OtelConfig { /** Whether to export telemetry. */ enable: boolean - /** Session id to report. A random one is generated when omitted. */ + /** Session ID to report. A random one is generated when omitted. */ sessionId?: string /** Metric export interval in milliseconds. Omit or pass 0 for the SDK default of 60000. */ exportIntervalMs?: number } -/** Enhancement parameters, all changeable while audio is being processed. */ +/** Configurable speech enhancement parameters. Values can be changed during processing. */ export declare const enum ProcessorParameter { /** - * Bypasses processing while preserving the algorithmic delay, so enhancement can be - * toggled without clicks or timing shifts. + * Bypasses enhancement while preserving the processing delay. * - * Range 0.0 - 1.0, where 0.0 is enhancement active and 1.0 is latency-compensated - * passthrough. Defaults to 0.0. + * This allows enhancement to be enabled or disabled without clicks or timing changes. + * + * Range: 0.0 to 1.0. At 0.0 enhancement is active; at 1.0 audio passes through with + * latency compensation. Default: 0.0. */ Bypass = 0, /** - * Tunes enhancement strength for a given STT engine, environment or UX requirement. + * Controls enhancement strength. + * + * Quail models apply stronger noise suppression at higher values, including suppression + * of competing speech for Voice Focus models. Rook models adjust the mix of original + * and enhanced audio. * - * Quail models suppress noise more aggressively as this rises (and, with Voice Focus, - * competing speech too); Rook models change the mixback. Range 0.0 - 1.0. + * Range: 0.0 to 1.0. */ EnhancementLevel = 1 } -/** Voice activity detection parameters, all changeable while audio is being processed. */ +/** Configurable voice activity detection parameters. Values can be changed during processing. */ export declare const enum VadParameter { /** - * How long the VAD keeps reporting speech after speech stops, which stabilizes - * detected -> not-detected transitions. + * Controls how long the VAD continues reporting speech after speech stops. * - * Speech is reported when at least half the blocks in the last - * `speechHoldDuration * 2` seconds contained speech, so ongoing speech extends the - * hold. Rounded to the model's window length, so reads may differ from writes. + * Speech is reported if at least half the blocks processed in the last + * `speechHoldDuration * 2` seconds contained speech. Additional speech during this + * period extends the detection period. * - * Range 0.0 to 300x the model window length, in seconds. Model-specific default. + * The duration is rounded to the nearest model window length, so the value read back + * may differ from the value set. + * + * Range: 0.0 to 300 times the model window length, in seconds. Default: model-specific. */ SpeechHoldDuration = 0, /** - * Probability threshold above which a block counts as speech, stabilizing how - * readily speech is detected at all. + * Sets the probability threshold for detecting speech in an audio block. + * + * A model probability above this threshold counts as speech. * - * Range 0.0 - 1.0. Model-specific default. + * Range: 0.0 to 1.0. Default: model-specific. */ Sensitivity = 1, /** - * How long speech must be present before the VAD reports it, which stabilizes - * not-detected -> detected transitions. + * Controls how long speech must be present before the VAD reports speech. + * + * The duration is rounded to the nearest model window length, so the value read back + * may differ from the value set. * - * Rounded to the model's window length, so reads may differ from writes. - * Range 0.0 - 1.0, in seconds. Model-specific default. + * Range: 0.0 to 1.0 seconds. Default: model-specific. */ MinimumSpeechDuration = 2 } diff --git a/src/analyzer.rs b/src/analyzer.rs index 62b2c06..f5f933d 100644 --- a/src/analyzer.rs +++ b/src/analyzer.rs @@ -15,26 +15,26 @@ use napi::{ use napi_derive::napi; use std::sync::{Arc, Mutex}; -/// Scores produced by {@link Analyzer#analyze}. +/// Results of analyzing an audio signal with {@link Analyzer}. /// -/// Every score runs 0.0 - 1.0. For all of them except `speakerLoudness`, lower means less -/// problematic audio. +/// Scores range from 0.0 to 1.0. For every field except `speakerLoudness`, lower values +/// indicate less problematic audio. #[napi(object)] pub struct AnalysisResult { - /// Headline score: how likely this audio is to break downstream models such as - /// speech-to-text, VAD, turn-taking or speech-to-speech. + /// Predicts the likelihood of failure in downstream models, including speech-to-text, + /// voice activity detection, turn-taking and speech-to-speech models. pub risk_score: f64, - /// How distant and reverberant the speaker sounds. + /// Measure of speaker distance and reverberation. pub speaker_reverb: f64, - /// How loud the speaker is. + /// Measure of speaker loudness. pub speaker_loudness: f64, - /// How much speech from people other than the main speaker is present. + /// Measure of interfering speech from sources other than the main speaker. pub interfering_speech: f64, - /// How much ambient or environmental noise is present. + /// Measure of ambient or environmental noise. pub noise: f64, - /// Artifacts from lossy speech codecs, e.g. a low bitrate or narrowband codec. + /// Measure of artifacts from lossy speech codecs, such as low bitrate or narrowband codecs. pub codec_degradation: f64, - /// Dropouts and discontinuities, e.g. from packet loss, frame erasure, jitter or CPU + /// Measure of audio dropouts and discontinuities from packet loss, frame erasure, jitter or CPU /// overload. pub packet_loss: f64, } @@ -53,39 +53,29 @@ impl From for AnalysisResult { } } -/// Analyzer for analysis models such as Tyto. +/// Analyzes audio quality using an analysis model, such as Tyto. /// -/// Buffering and analysis are separate calls: {@link Analyzer#buffer} is cheap enough for -/// the audio path, while running the model is not. Analysis comes in two forms, -/// {@link Analyzer#analyzeAsync} on a worker thread and {@link Analyzer#analyze} on the -/// calling thread. +/// Call {@link Analyzer#initialize}, then {@link Analyzer#buffer} to collect mono audio. +/// The model determines how much audio is retained; older samples are discarded as new +/// audio arrives. /// -/// Only a fixed span of audio is retained, determined by the model; older audio is -/// discarded as more is buffered. -/// -/// The SDK splits this into a collector and an analyzer so the two halves can live on -/// different threads. A class instance cannot cross into a Node worker, so both are -/// exposed as one object here, and the split shows through: {@link Analyzer#buffer} drives -/// the collector on the calling thread, while {@link Analyzer#analyzeAsync} moves the -/// analyzer half onto a worker. The SDK guarantees the two are safe to use concurrently. +/// Run {@link Analyzer#analyzeAsync} to analyze the buffered audio on a libuv worker +/// thread, or {@link Analyzer#analyze} to run it on the calling thread. Analysis is +/// computationally expensive and should not run in audio processing callbacks. +/// Buffering remains synchronous and can continue while async analysis is running. #[napi(custom_finalize)] pub struct Analyzer { - // Neither half borrows the other, nor the model: `'static` here is the model weights' - // lifetime, which `Model.fromFile` satisfies by memory-mapping the file. - // - // Only the analyzer half is shared, so `buffer`, the one call on the audio path, takes - // no lock and cannot contend with an analysis running on a worker thread. + // Both SDK objects own their state and retain the memory-mapped model as needed. + // Only the analyzer is shared with workers; collector access does not take its lock. collector: DisposableSlot, analyzer: Arc>>>, } impl ObjectFinalize for Analyzer { fn finalize(mut self, env: Env) -> Result<()> { - // An in-flight `AnalyzeTask` holds the analyzer `Arc`, so that half is dropped only - // after the worker finishes, and its footprint report is left to whoever holds the - // last handle. The collector drops here, possibly mid-analysis, which - // `aic_collector_destroy` permits: the paired halves are destroyed independently, in - // any order. Both releases are no-ops if `dispose()` got there first. + // A pending task can retain the analyzer after this finalizer; see `DisposableSlot`. + // The SDK permits the collector and analyzer to be destroyed independently, so the + // collector can be released during analysis. Both releases are idempotent. self.collector.release(env); if Arc::strong_count(&self.analyzer) == 1 { lock(&self.analyzer).release(env); @@ -96,7 +86,12 @@ impl ObjectFinalize for Analyzer { #[napi] impl Analyzer { - /// Creates an analyzer from an analysis model. Other model types are rejected. + /// Creates an analyzer from an analysis model. + /// + /// Construction allocates memory. Call {@link Analyzer#initialize} before buffering audio. + /// + /// @param model - Analysis model. Other model types are rejected. + /// @param licenseKey - SDK license key from . #[napi(constructor)] pub fn new(env: Env, model: &Model, license_key: String) -> Result { let model_inner = model.inner()?; @@ -114,25 +109,29 @@ impl Analyzer { }) } - /// Destroys the native collector and analyzer immediately, releasing their memory - /// without waiting for garbage collection. + /// Destroys the native collector and analyzer without waiting for garbage collection. /// - /// Every later method throws; calling `dispose()` again does nothing. Blocks until an - /// in-flight `analyzeAsync` on a worker thread finishes. + /// After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. + /// This call blocks the calling thread while a worker holds the analyzer lock. + /// Queued analysis that acquires the lock after disposal rejects its promise. #[napi] pub fn dispose(&mut self, env: Env) { - // The collector is released without taking the analyzer lock, so it can be destroyed - // while an `analyzeAsync` is in flight; see the finalizer for why that is safe. The - // analyzer half is released under the lock, so this call blocks until that analysis - // finishes. Both releases are idempotent. + // The SDK allows collector destruction during analysis. Releasing the analyzer + // requires its lock and waits if a worker currently holds it. self.collector.release(env); lock(&self.analyzer).release(env); } - /// Configures the analyzer for an audio format. Must be called before buffering. + /// Configures the analyzer's audio collector. /// - /// The model's optimal sample rate and block size avoid internal resampling and - /// rebuffering. Allocates, so keep it off the audio path. + /// Call this before buffering audio. Use {@link Model#getOptimalSampleRate} and + /// {@link Model#getOptimalBlockSize} to avoid internal resampling and rebuffering. + /// This method allocates memory; avoid calling it from audio processing callbacks. + /// + /// @param sampleRate - Audio sample rate in Hz. + /// @param blockSize - Number of mono samples per block. + /// @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + /// Blocks larger than `blockSize` are always rejected. #[napi] pub fn initialize( &mut self, @@ -147,37 +146,36 @@ impl Analyzer { ))) } - /// Buffers a mono audio block for later analysis, leaving the audio unmodified. + /// Buffers a mono audio block for later analysis without modifying the input. + /// + /// Call {@link Analyzer#initialize} first. The block must contain exactly `blockSize` + /// samples, or at most `blockSize` if `variableBlockSize` is enabled. + /// Buffering does not acquire the analyzer lock and can run during async analysis. #[napi] pub fn buffer(&mut self, audio: Float32Array) -> Result<()> { map_err(self.collector.get_mut()?.buffer(&audio)) } - /// Runs the analysis model over the buffered audio, on the calling thread. + /// Analyzes buffered audio on the calling thread and returns the scores. /// - /// The model consumes a fixed span of audio. Calling this before that much has been - /// buffered analyzes what is there, padded with silence. + /// The model analyzes a fixed duration of audio. If less audio has been buffered, the + /// remaining input is padded with silence. /// - /// Analysis is mono. Mix multichannel audio down, or use one analyzer per channel. - /// - /// This call is expensive and blocks. Prefer {@link Analyzer#analyzeAsync} unless - /// nothing else is waiting on the event loop. + /// This method blocks the calling thread. Use {@link Analyzer#analyzeAsync} to keep the + /// event loop available. For multichannel audio, mix down to mono or use one analyzer + /// per channel. #[napi] pub fn analyze(&self) -> Result { map_err(lock(&self.analyzer).get_mut()?.analyze_buffered()).map(AnalysisResult::from) } - /// Runs the analysis model over the buffered audio on a worker thread. + /// Analyzes buffered audio on a libuv worker thread and returns a promise for the scores. /// - /// Same result as {@link Analyzer#analyze}, off the event loop. Prefer this wherever - /// the analysis model is too expensive to run on the calling thread. + /// Uses the same analysis and silence padding as {@link Analyzer#analyze}. + /// {@link Analyzer#buffer} can continue collecting audio while analysis runs. /// - /// {@link Analyzer#buffer} stays synchronous and takes no lock, so audio can keep - /// arriving while an analysis is in flight; the SDK guarantees the collector and - /// analyzer halves are safe to use concurrently. The other methods here do take the - /// analyzer's lock, so calling {@link Analyzer#analyze}, {@link Analyzer#reset} or - /// {@link Analyzer#terminateSession} while this is pending blocks the calling thread - /// until it finishes. + /// Synchronous calls to `analyze`, `reset`, `updateBearerToken`, `terminateSession` and + /// `dispose` wait for the analyzer lock and may block while analysis is running. #[napi(ts_return_type = "Promise")] pub fn analyze_async(&self) -> AsyncTask { AsyncTask::new(AnalyzeTask { @@ -185,30 +183,41 @@ impl Analyzer { }) } - /// Clears buffered audio and internal state, keeping the configured audio settings. + /// Clears buffered audio and internal state while preserving the configured audio settings. #[napi] pub fn reset(&self) -> Result<()> { map_err(lock(&self.analyzer).get_mut()?.reset()) } - /// Swaps in a renewed JWT without tearing down the analyzer. + /// Replaces the bearer token on the running analyzer. + /// + /// Use this to refresh a JWT without recreating the instance. Both the original license + /// key and the new token must be JWTs. If this call fails, the previous token remains active. /// - /// Only works when both the original key and the new token are JWTs. On failure the - /// call is a no-op and the previous token stays active. + /// A successful call validates the token's format and applies it immediately. Backend + /// acceptance is checked later. If the backend rejects the token, the SDK retries with + /// backoff; analysis may be rejected if no accepted token arrives in time. + /// Supply a valid token to recover the session. + /// + /// This method allocates memory and takes a mutex. Avoid calling it from audio processing callbacks. #[napi] pub fn update_bearer_token(&self, token: String) -> Result<()> { map_err(lock(&self.analyzer).get_mut()?.update_bearer_token(&token)) } - /// Ends this analyzer's telemetry session, after which it can no longer analyze audio. + /// Terminates the telemetry session associated with this analyzer. + /// + /// Once termination is handled, the analyzer can no longer analyze buffered audio. + /// The session also ends when the native analyzer is destroyed. This method may block; + /// avoid calling it from audio processing callbacks. If another session is still active, + /// termination can complete asynchronously. #[napi] pub fn terminate_session(&self) -> Result<()> { map_err(lock(&self.analyzer).get_mut()?.terminate_session()) } } -/// Holds only the analyzer half, so the collector stays on the JS thread where `buffer` -/// can keep reaching it while this runs. +/// Runs analysis on a worker while the collector remains accessible on the JavaScript thread. pub struct AnalyzeTask { analyzer: Arc>>>, } diff --git a/src/disposable_slot.rs b/src/disposable_slot.rs index 7046d76..3b0aec5 100644 --- a/src/disposable_slot.rs +++ b/src/disposable_slot.rs @@ -1,10 +1,7 @@ -//! One native SDK object, destroyed exactly once, with its footprint reported to V8 for -//! as long as it lives. +//! Owns a native SDK object and tracks its estimated memory usage in V8. //! -//! Every binding class holds its SDK object in a slot, so the disposed error and the -//! footprint reporting live here instead of in each class. The slot has no interior -//! mutability; the classes that share their object with tasks on the libuv pool wrap it -//! in an `Arc>` and reach it through [`lock`]. +//! All binding classes use this slot for disposal checks and memory reporting. +//! Classes shared with libuv tasks wrap it in `Arc>`. use std::sync::{Mutex, MutexGuard}; @@ -24,7 +21,7 @@ use crate::{ /// /// A slot dropped without `release` having run still destroys the object, but its bytes /// stay reported: withdrawing them needs an `Env`, which `Drop` does not have. This -/// happens when the last JS handle onto a shared object is finalized while a task still +/// happens when the last JS handle to a shared object is finalized while a task still /// holds a clone, and leaves V8 over-reported for the rest of the process. pub(crate) struct DisposableSlot { inner: Option, @@ -71,12 +68,10 @@ impl DisposableSlot { } } -/// Locks a shared slot, recovering the guard if the lock is poisoned. +/// Locks a shared slot, recovering the guard if the mutex is poisoned. /// -/// A blocking `Mutex` suits the callers: the tasks that share a slot run on libuv -/// workers. Poisoning would take a panic inside an SDK call, which leaves the slot's own -/// `Option` intact, so recovering the guard is safe and keeps later calls working, -/// disposal included. +/// Recovery allows disposal after a task panics. The slot's `Option` still records +/// whether the native object has been released. pub(crate) fn lock(slot: &Mutex) -> MutexGuard<'_, T> { slot.lock().unwrap_or_else(|poisoned| poisoned.into_inner()) } diff --git a/src/error.rs b/src/error.rs index 1effdc1..2b805dd 100644 --- a/src/error.rs +++ b/src/error.rs @@ -1,9 +1,9 @@ use aic_sdk::AicError; -/// Newtype so the SDK's error can be converted into a JS exception. +/// Local wrapper for converting an SDK error into a JavaScript exception. /// -/// `AicError` and `napi::Error` are both foreign to this crate, so the -/// conversion needs a local type to hang the impl on. +/// Rust's orphan rules require a local type because both `AicError` and `napi::Error` +/// are defined in dependencies. pub struct JsAicError(pub AicError); impl From for JsAicError { @@ -14,12 +14,7 @@ impl From for JsAicError { impl From for napi::Error { fn from(error: JsAicError) -> Self { - // `AicError`'s `Display` messages are user-facing and already explain how to - // recover (see `aic-sdk`'s error.rs), so this passes them through verbatim. - // - // Every variant maps to a plain `Error`. Mapping `ParameterOutOfRange` to a - // `RangeError`, as the wasm binding does, would need a distinct return type per - // method in napi-rs. + // Preserve the SDK's error message. All SDK errors become plain JavaScript Errors. napi::Error::new(napi::Status::GenericFailure, error.0.to_string()) } } diff --git a/src/lib.rs b/src/lib.rs index d8aaa06..ecda4b9 100644 --- a/src/lib.rs +++ b/src/lib.rs @@ -2,8 +2,8 @@ //! Node.js bindings for the ai-coustics SDK. //! -//! Thin napi-rs layer over the `aic-sdk` crate: the classes here own the corresponding SDK -//! types directly, so lifetimes, cleanup and thread-safety are handled upstream. +//! The napi-rs classes wrap the Rust SDK's model, processing, VAD and analysis types. +//! Native objects are released through explicit disposal or JavaScript finalization. use napi_derive::napi; @@ -24,28 +24,23 @@ pub use processor_async::*; pub use vad::*; pub use vad_async::*; -/// Telemetry id assigned to this binding. Must match `SdkWrapper::NodeJs` in +/// Telemetry ID assigned to this binding. Must match `SdkWrapper::NodeJs` in /// `aic-sdk-telemetry`. const SDK_WRAPPER_ID_NODE: u32 = 4; -/// Claims the Node telemetry id, unless a wrapper embedding this package claimed its own -/// first via [`set_sdk_id`]. +/// Sets the Node telemetry ID unless an embedding wrapper has already set one. /// -/// The id lives in a `OnceLock` upstream: the first write wins and every later write is -/// silently discarded. `aic-sdk` sets `2` ("Rust") inside each of its `Processor`, `Vad` -/// and analyzer constructors, so every constructor here claims `4` before delegating. -/// -/// Not called at module load: an embedder can only reach [`set_sdk_id`] once the module -/// is loaded, so the window between load and first construction stays open for their -/// write. +/// The upstream `OnceLock` accepts only the first write. Call this before delegating to +/// an SDK constructor, which would otherwise set the Rust wrapper ID (2). +/// Delaying this until construction lets embedders call [`set_sdk_id`] after module load. pub(crate) fn claim_sdk_id() { - // SAFETY: `4` is the wrapper id assigned to this binding by ai-coustics. + // SAFETY: `4` is the wrapper ID assigned to this binding by ai-coustics. unsafe { aic_sdk::set_sdk_id(SDK_WRAPPER_ID_NODE) }; } -/// Overrides the telemetry wrapper id. Internal only, for ai-coustics wrappers embedding +/// Overrides the telemetry wrapper ID. Internal only, for ai-coustics wrappers embedding /// this package (e.g. the LiveKit plugin): call before constructing any `Processor`, `Vad` -/// or `Analyzer`, whose constructors otherwise claim the id for this SDK. The id can only +/// or `Analyzer`, whose constructors otherwise claim the ID for this SDK. The ID can only /// be set once per process; later writes are silently discarded. #[napi(js_name = "_setSdkId")] pub fn set_sdk_id(id: u32) { @@ -53,17 +48,15 @@ pub fn set_sdk_id(id: u32) { unsafe { aic_sdk::set_sdk_id(id) }; } -/// The version of the underlying native SDK, e.g. `"0.23.0"`. +/// Returns the version of the underlying native SDK. /// -/// Not necessarily this package's version. +/// This may differ from the Node.js package version. #[napi] pub fn get_version() -> String { aic_sdk::get_sdk_version().to_owned() } -/// The model file format version this SDK can load. -/// -/// Model URLs are versioned by this number, so it decides which model files are usable. +/// Returns the model file format version supported by this SDK. #[napi] pub fn get_compatible_model_version() -> u32 { aic_sdk::get_compatible_model_version() diff --git a/src/mem.rs b/src/mem.rs index 75d844a..66dfd6d 100644 --- a/src/mem.rs +++ b/src/mem.rs @@ -1,18 +1,11 @@ -//! Reports the native footprint of SDK objects to V8's garbage collector. +//! Reports estimated native memory usage to V8's garbage collector. //! -//! Each binding class is a small JS object in front of a much larger native allocation, -//! from ~200 KiB for a processor up to the size of the model weights. V8 only sees the JS -//! side, so without a hint it has no reason to collect dropped instances and a workload -//! that creates processors per unit of work grows unchecked. -//! `Env::adjust_external_memory` (`napi_adjust_external_memory`) reports that hidden cost. +//! V8 cannot infer native allocations from the small JavaScript wrapper objects. +//! `DisposableSlot` adds an estimate on construction and subtracts it on explicit +//! release or finalization. This helps GC account for models and processing state. //! -//! [`DisposableSlot`](crate::disposable_slot::DisposableSlot) does the reporting. It adds -//! its object's footprint on construction and withdraws it again on release, whether that -//! comes from `dispose()` or from the class finalizer, whichever gets there first. -//! -//! The SDK exposes no per-instance memory query, so the footprints below are per-class -//! constants: measured estimates, rounded up. Over-reporting only costs some extra GC -//! work; under-reporting would let the growth back in. +//! The SDK has no per-instance memory query. The constants below are rounded-up +//! measurements; model estimates use the file size. use std::path::Path; @@ -38,7 +31,7 @@ pub(crate) const PROCESSOR_BYTES: i64 = 512 * KIB; /// models. /// /// Reported separately from [`COLLECTOR_BYTES`] because the two halves are destroyed -/// independently: the collector can go while a worker thread still analyzes, so a single +/// independently: the collector can be destroyed during analysis, so a single /// report for the pair would be released too early. pub(crate) const ANALYZER_BYTES: i64 = 14 * MIB; @@ -47,7 +40,7 @@ pub(crate) const ANALYZER_BYTES: i64 = 14 * MIB; /// 5 s span, so 2 MiB leaves ~3x headroom. pub(crate) const COLLECTOR_BYTES: i64 = 2 * MIB; -/// Fallback footprint for a `Model` when its file cannot be stat'd. The weights are +/// Fallback footprint for a `Model` when file metadata is unavailable. The weights are /// memory-mapped, so the resident share approaches the file size as pages are touched. const MODEL_FALLBACK_BYTES: i64 = 64 * MIB; @@ -64,7 +57,7 @@ pub(crate) fn adjust(env: Env, delta_bytes: i64) { } /// The footprint to report for a model loaded from `path`. The weights are memory-mapped, -/// so this is the file size, or [`MODEL_FALLBACK_BYTES`] when the file cannot be stat'd. +/// so this is the file size, or [`MODEL_FALLBACK_BYTES`] when file metadata is unavailable. pub(crate) fn model_bytes(path: &Path) -> i64 { std::fs::metadata(path) .map(|meta| meta.len() as i64) diff --git a/src/model.rs b/src/model.rs index 53d4ff1..d63b01c 100644 --- a/src/model.rs +++ b/src/model.rs @@ -9,16 +9,13 @@ use napi_derive::napi; /// A loaded ai-coustics model. /// -/// One model can back multiple processors, VADs and analyzers, according to its type. -/// The underlying model data is kept alive by every object created from it through -/// internal reference counting, so this handle may be released first. +/// The same model can be used to create multiple independent instances of a compatible +/// processor, VAD or analyzer. Each instance retains a reference to the model data, +/// so the `Model` handle can be disposed or garbage collected first. #[napi(custom_finalize)] pub struct Model { - // `from_file` memory-maps the file instead of borrowing a caller-owned buffer, so the - // SDK model is `'static` and needs no lifetime plumbing here. - // - // The slot's footprint is the mmap'd weights, so unlike the other classes it is a - // per-instance value: the model file's size. + // The memory-mapped model owns its data and has a `'static` lifetime. + // Report the file size as this instance's native memory estimate. slot: DisposableSlot>, } @@ -41,12 +38,14 @@ impl Model { impl Model { /// Loads a model from a `.aicmodel` file. /// - /// The model data is memory-mapped, not copied into the process, so the file must - /// not be modified or deleted while this model, or any object created from it, - /// is alive. + /// The SDK memory-maps the file. Do not modify or delete it while the model or any + /// processor, VAD or analyzer created from it is still alive. /// - /// Browse available models at , or fetch one with - /// {@link Model.download}. + /// Download models with {@link Model.download}. Available model IDs are listed at + /// . + /// + /// @param path - Path to the model file. + /// @throws If the file cannot be loaded or its format is incompatible with this SDK. #[napi(factory)] pub fn from_file(env: Env, path: String) -> Result { let inner = map_err(aic_sdk::Model::from_file(&path))?; @@ -57,24 +56,26 @@ impl Model { }) } - /// Unmaps the model file immediately, releasing its footprint without waiting for - /// garbage collection. + /// Releases this handle's reference to the native model without waiting for garbage collection. + /// + /// Instances created from this model remain usable. The file mapping is released when + /// its last reference is destroyed. /// - /// Objects already created from the model keep working: the SDK keeps the weights - /// alive through internal reference counting. Every later method on this handle - /// throws; calling `dispose()` again does nothing. + /// After disposal, all methods except `dispose()` throw. Repeated disposal has no effect. #[napi] pub fn dispose(&mut self, env: Env) { self.slot.release(env); } - /// Downloads a model from the ai-coustics artifact CDN and resolves to its path. + /// Downloads a model from the ai-coustics CDN and returns a promise for its file path. /// - /// The manifest is re-fetched on every call so the newest compatible model version - /// is always used. An existing file with a matching checksum is not re-downloaded; - /// one with a mismatching checksum is replaced. + /// Each call fetches the manifest to select the latest compatible model version. + /// An existing file is reused if its checksum matches; otherwise it is replaced. + /// The download runs on Node's libuv thread pool. /// - /// The download runs on a worker thread, so it does not block the event loop. + /// @param modelId - Model ID listed at . + /// @param downloadDir - Directory in which to store the model. + /// @returns A promise for the downloaded or cached model's path. // napi cannot infer an `AsyncTask`'s resolved type; without the annotation the // generated d.ts says `Promise`. #[napi(ts_return_type = "Promise")] @@ -85,31 +86,30 @@ impl Model { }) } - /// The model identifier, e.g. `quail-vf-2.2-s-16khz`. + /// Returns the model identifier, including its build and format version suffixes. #[napi] pub fn get_id(&self) -> Result { Ok(self.inner()?.id().to_owned()) } - /// The sample rate in Hz the model was trained for. + /// Returns the sample rate in Hz for which the model was trained. /// - /// Audio at any rate can be processed, but a model only enhances frequencies up to - /// its own Nyquist limit, so matching this rate gives the best quality. + /// The SDK resamples audio at other supported rates internally. Enhancement is limited + /// to frequencies below half the model's sample rate. #[napi] pub fn get_optimal_sample_rate(&self) -> Result { Ok(self.inner()?.optimal_sample_rate()) } - /// The block size that avoids internal buffering at `sampleRate`. + /// Returns the optimal block size in samples for the given sample rate. + /// + /// Use this value to avoid additional buffering latency. The block size depends on the + /// sample rate because each model processes a fixed duration of audio. /// - /// Any other block size adds buffering latency on top of the base processing delay. - /// The value changes with the sample rate, because the model works on a fixed time - /// window: a 10 ms window is 480 samples at 48 kHz but 160 at 16 kHz. + /// @param sampleRate - Audio sample rate in Hz. #[napi] pub fn get_optimal_block_size(&self, sample_rate: u32) -> Result { - // The SDK reports sizes as `usize`, which napi would marshal as a JS BigInt, and a - // BigInt block size throws on `new Float32Array(n)` and on arithmetic against plain - // numbers. Block sizes are a few thousand samples at most, so u32 is ample. + // Use `u32` so block sizes are exposed as JavaScript numbers for typed-array lengths. Ok(self.inner()?.optimal_block_size(sample_rate) as u32) } } diff --git a/src/processor.rs b/src/processor.rs index 0b64858..4e1967e 100644 --- a/src/processor.rs +++ b/src/processor.rs @@ -9,19 +9,23 @@ use crate::{ use napi::{Env, bindgen_prelude::Float32Array, bindgen_prelude::ObjectFinalize}; use napi_derive::napi; -/// Enhancement parameters, all changeable while audio is being processed. +/// Configurable speech enhancement parameters. Values can be changed during processing. #[napi] pub enum ProcessorParameter { - /// Bypasses processing while preserving the algorithmic delay, so enhancement can be - /// toggled without clicks or timing shifts. + /// Bypasses enhancement while preserving the processing delay. /// - /// Range 0.0 - 1.0, where 0.0 is enhancement active and 1.0 is latency-compensated - /// passthrough. Defaults to 0.0. + /// This allows enhancement to be enabled or disabled without clicks or timing changes. + /// + /// Range: 0.0 to 1.0. At 0.0 enhancement is active; at 1.0 audio passes through with + /// latency compensation. Default: 0.0. Bypass = 0, - /// Tunes enhancement strength for a given STT engine, environment or UX requirement. + /// Controls enhancement strength. + /// + /// Quail models apply stronger noise suppression at higher values, including suppression + /// of competing speech for Voice Focus models. Rook models adjust the mix of original + /// and enhanced audio. /// - /// Quail models suppress noise more aggressively as this rises (and, with Voice Focus, - /// competing speech too); Rook models change the mixback. Range 0.0 - 1.0. + /// Range: 0.0 to 1.0. EnhancementLevel = 1, } @@ -34,15 +38,15 @@ impl From for aic_sdk::ProcessorParameter { } } -/// Per-instance OpenTelemetry settings. +/// OpenTelemetry configuration for a processor or VAD instance. /// -/// Overrides the environment-based defaults (e.g. `AIC_SDK_OTEL_ENABLE`) for the one -/// processor or VAD it is passed to. +/// Overrides the SDK's environment-based telemetry settings, such as +/// `AIC_SDK_OTEL_ENABLE`, for the instance receiving this configuration. #[napi(object)] pub struct OtelConfig { /// Whether to export telemetry. pub enable: bool, - /// Session id to report. A random one is generated when omitted. + /// Session ID to report. A random one is generated when omitted. pub session_id: Option, /// Metric export interval in milliseconds. Omit or pass 0 for the SDK default of 60000. pub export_interval_ms: Option, @@ -58,10 +62,7 @@ impl From for aic_sdk::OtelConfig { } } -/// Builds the SDK audio config shared by processors, VADs and analyzers. -/// -/// Block sizes cross the JS boundary as `u32`; the SDK's `usize` would reach JS as a -/// BigInt. +/// Converts JavaScript audio settings into the SDK configuration. Block sizes use JS numbers. pub(crate) fn audio_config( sample_rate: u32, block_size: u32, @@ -74,16 +75,16 @@ pub(crate) fn audio_config( } } -/// Speech enhancement processor. +/// Processes mono audio with an enhancement or bypass model. /// -/// Built from an enhancement or bypass model. Use {@link Vad} for dedicated VAD models -/// and {@link Analyzer} for analysis models; passing the wrong kind throws. +/// Call {@link Processor#initialize} before processing audio. Each processor maintains +/// its own state; create one instance per audio stream. Multiple processors can share +/// the same {@link Model}. /// -/// Create several processors to handle multiple streams or to switch models at runtime. +/// Use {@link ProcessorAsync} to run processing on Node's libuv thread pool. #[napi(custom_finalize)] pub struct Processor { - // No lock: every method here runs on the JS thread. Only the async class shares its - // slot with tasks on the libuv pool. + // Only the JavaScript thread accesses this slot. Async classes use a shared, locked slot. slot: DisposableSlot>, } @@ -97,10 +98,15 @@ impl ObjectFinalize for Processor { #[napi] impl Processor { - /// Creates a processor from an enhancement or bypass model. + /// Creates a new speech enhancement processor. + /// + /// Construction is synchronous and throws if creation fails. Call + /// {@link Processor#initialize} before processing audio. /// - /// Telemetry follows the runtime environment; pass `otelConfig` to override it for this - /// instance. + /// @param model - Enhancement or bypass model. Other model types are rejected. + /// @param licenseKey - SDK license key from . + /// @param otelConfig - Optional telemetry configuration. When omitted, telemetry follows + /// the runtime environment. #[napi(constructor)] pub fn new( env: Env, @@ -123,22 +129,25 @@ impl Processor { }) } - /// Destroys the native processor immediately, releasing its memory and telemetry - /// session without waiting for garbage collection. + /// Destroys the native processor and releases its telemetry session. /// - /// Every later method throws; calling `dispose()` again does nothing. + /// Use this for cleanup at a specific point instead of waiting for garbage collection. + /// After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. #[napi] pub fn dispose(&mut self, env: Env) { self.slot.release(env); } - /// Configures the processor for an audio format. Must be called before processing. + /// Configures the processor for the given audio format. /// - /// For the lowest delay use {@link Model#getOptimalSampleRate} and - /// {@link Model#getOptimalBlockSize}. Allocates, so keep it off the audio path. + /// Call this method before processing audio. Use {@link Model#getOptimalSampleRate} and + /// {@link Model#getOptimalBlockSize} for the lowest delay. + /// This method allocates memory; avoid calling it from audio processing callbacks. /// - /// With `variableBlockSize` enabled (default `false`), calls shorter than `blockSize` - /// are permitted at the cost of extra delay; longer calls are always rejected. + /// @param sampleRate - Audio sample rate in Hz. + /// @param blockSize - Number of mono samples per block. + /// @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + /// Variable block sizes can add buffering latency. Larger blocks are always rejected. #[napi] pub fn initialize( &mut self, @@ -155,18 +164,17 @@ impl Processor { /// Enhances a mono audio block in place. /// - /// The block must be exactly `blockSize` samples, or at most `blockSize` if - /// `variableBlockSize` was enabled. + /// Call {@link Processor#initialize} first. The block must contain exactly `blockSize` + /// samples, or at most `blockSize` if `variableBlockSize` is enabled. + /// If the input uses a SharedArrayBuffer, prevent other workers from accessing it during + /// this call. #[napi] pub fn process(&mut self, mut audio: Float32Array) -> Result<()> { - // Taken by value without copying: `Float32Array` is a view onto the caller's - // ArrayBuffer, so the writes below land in the JS-owned buffer. + // `Float32Array` references the caller's ArrayBuffer; processing writes into it directly. // - // SAFETY: `as_mut` is unsafe because JS could mutate the backing ArrayBuffer - // concurrently. It cannot here: the call is synchronous, so no JS runs while the - // slice is alive, the slice never escapes this function, and each Node thread has - // its own isolate. The exception is a SharedArrayBuffer written by another worker - // mid-call, which no in-place API can guard against. + // SAFETY: The call is synchronous and the mutable slice does not escape this function. + // JavaScript in this isolate cannot access the buffer during the call. A caller using + // SharedArrayBuffer must prevent concurrent access from other workers. let samples = unsafe { audio.as_mut() }; map_err(self.slot.get_mut()?.process(samples)) @@ -174,7 +182,7 @@ impl Processor { /// Creates a handle for reading and writing this processor's parameters and state. /// - /// Each call returns an independent handle onto the same processor. + /// Each call returns an independent handle to the same processor. #[napi] pub fn get_context(&self) -> Result { Ok(ProcessorContext { @@ -182,11 +190,14 @@ impl Processor { }) } - /// Ends this processor's telemetry session, after which it can no longer process audio. + /// Terminates the telemetry session associated with this processor. /// - /// A session is closed automatically when the processor is collected, but GC timing is - /// not guaranteed, so call this on a lifecycle event instead. May block, so keep it off - /// the audio path. + /// Once termination is handled, the processor can no longer process audio. + /// The session also ends when the native object is destroyed. Use this method when + /// termination must be requested at a specific lifecycle event. + /// + /// This method may block. Avoid calling it from audio processing callbacks. + /// If another session is still active, termination can complete asynchronously. #[napi] pub fn terminate_session(&mut self) -> Result<()> { map_err(self.slot.get_mut()?.terminate_session()) @@ -212,13 +223,13 @@ impl ProcessorContext { map_err(self.inner.set_parameter(parameter.into(), value as f32)) } - /// Reads the current value of a parameter. + /// Returns the current value of an enhancement parameter. #[napi] pub fn get_parameter(&self, parameter: ProcessorParameter) -> Result { map_err(self.inner.parameter(parameter.into())).map(f64::from) } - /// Total delay the processor applies to the audio, in samples at the initialized rate. + /// Returns the audio delay in samples at the configured sample rate. /// /// Covers algorithmic delay plus any buffering from a non-optimal block size. Before /// initialization it reports the base delay at the model's optimal settings. @@ -227,20 +238,26 @@ impl ProcessorContext { self.inner.audio_delay() as u32 } - /// Clears internal state and buffers, keeping the configured audio settings. + /// Clears internal state and buffers while preserving the configured audio settings. /// - /// Call this on a stream discontinuity or when seeking, to keep earlier audio from - /// bleeding into the output. + /// Call this when the stream is interrupted or when seeking to prevent previous audio + /// from affecting the output. #[napi] pub fn reset(&self) -> Result<()> { map_err(self.inner.reset()) } - /// Swaps in a renewed JWT without interrupting processing. + /// Replaces the bearer token on the running processor. + /// + /// Use this to refresh a JWT without recreating the instance. Both the original license + /// key and the new token must be JWTs. If this call fails, the previous token remains active. + /// + /// A successful call validates the token's format and applies it immediately. Backend + /// acceptance is checked later. If the backend rejects the token, the SDK retries with + /// backoff; processing is eventually disabled if no accepted token arrives in time. + /// Supply a valid token to recover the session. /// - /// Only works when both the original key and the new token are JWTs. On failure the - /// call is a no-op and the previous token stays active. On success the swap is applied - /// immediately and is not gated on backend acceptance. + /// This method allocates memory and takes a mutex. Avoid calling it from audio processing callbacks. #[napi] pub fn update_bearer_token(&self, token: String) -> Result<()> { map_err(self.inner.update_bearer_token(&token)) diff --git a/src/processor_async.rs b/src/processor_async.rs index 368f3b8..0e8d66e 100644 --- a/src/processor_async.rs +++ b/src/processor_async.rs @@ -14,24 +14,22 @@ use napi::{ use napi_derive::napi; use std::sync::{Arc, Mutex}; -/// Speech enhancement processor that keeps its work off the main thread. +/// Speech enhancement for use in async applications. /// -/// The same processing as {@link Processor}, but each call returns a promise and runs on -/// Node's libuv thread pool, so the event loop stays responsive. Prefer this when other -/// work shares that loop, as in a server; prefer {@link Processor} on a dedicated audio -/// thread or in a batch script, where nothing else needs the loop. +/// Initialization, processing and context creation run on Node's libuv thread pool and +/// return promises. Construction and disposal are synchronous. /// -/// Mirrors `ProcessorAsync` in the Rust SDK. +/// Use {@link Processor} when processing should run on the calling thread. /// -/// ### Concurrency +/// ### Threading /// -/// One instance handles one stream. Do not start a second {@link ProcessorAsync#process} -/// before the first resolves: libuv completes work items out of order, which would -/// desync the stream. To process several streams at once, create several instances. +/// Use one instance per stream and await each operation before submitting the next. +/// Concurrent calls on one instance are not guaranteed to execute in submission order. +/// Use separate instances to process multiple streams concurrently. /// -/// The libuv pool is four threads by default and is shared with `fs`, `dns` and `crypto`. -/// Raise `UV_THREADPOOL_SIZE` before Node starts to run more streams in parallel. -/// `AIC_NUM_THREADS` has no effect: it sizes a rayon pool this binding does not use. +/// The libuv pool defaults to four threads and is shared with filesystem, DNS and crypto +/// work. Set `UV_THREADPOOL_SIZE` before starting Node to change its size. +/// `AIC_NUM_THREADS` does not apply to these bindings. #[napi(custom_finalize)] pub struct ProcessorAsync { slot: Arc>>>, @@ -39,9 +37,9 @@ pub struct ProcessorAsync { impl ObjectFinalize for ProcessorAsync { fn finalize(self, env: Env) -> Result<()> { - // Only the last handle destroys the native object; while other handles or in-flight - // tasks hold an `Arc`, this leaves the object and its footprint report to them. - // A no-op if `dispose()` already ran. + // Release the object and V8 memory estimate only for the last shared handle. + // A pending task can retain the slot beyond finalization; see `DisposableSlot`. + // `release` has no effect if the object was already disposed. if Arc::strong_count(&self.slot) == 1 { lock(&self.slot).release(env); } @@ -51,13 +49,15 @@ impl ObjectFinalize for ProcessorAsync { #[napi] impl ProcessorAsync { - /// Creates a processor from an enhancement or bypass model. + /// Creates a new async speech enhancement processor. /// - /// Construction is synchronous and throws on failure, as in the Rust SDK; only the - /// audio work runs on a worker thread. + /// Construction is synchronous and throws if creation fails. Await + /// {@link ProcessorAsync#initialize} or {@link ProcessorAsync#withConfig} before processing audio. /// - /// Telemetry follows the runtime environment; pass `otelConfig` to override it for this - /// instance. + /// @param model - Enhancement or bypass model. Other model types are rejected. + /// @param licenseKey - SDK license key from . + /// @param otelConfig - Optional telemetry configuration. When omitted, telemetry follows + /// the runtime environment. #[napi(constructor)] pub fn new( env: Env, @@ -85,26 +85,26 @@ impl ProcessorAsync { }) } - /// Destroys the native processor immediately, releasing its memory and telemetry - /// session without waiting for garbage collection. + /// Destroys the native processor and releases its telemetry session. /// - /// Every later method throws; calling `dispose()` again does nothing. Blocks until - /// in-flight work on the libuv pool finishes. + /// Use this for cleanup at a specific point instead of waiting for garbage collection. + /// After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. + /// + /// This call blocks the calling thread while a worker holds the instance lock. + /// Queued work that acquires the lock after disposal rejects its promise. #[napi] pub fn dispose(&self, env: Env) { lock(&self.slot).release(env); } - /// Initializes the processor and resolves to a handle onto it, for chaining off the - /// constructor: + /// Initializes the processor and returns a promise for a handle to the initialized instance. /// - /// ```js - /// const processor = await new ProcessorAsync(model, licenseKey).withConfig(48000, 480) - /// ``` + /// Uses the same configuration as {@link ProcessorAsync#initialize}. The returned handle and + /// this object share the same native instance; disposing either invalidates both. /// - /// The handle it resolves to drives the same underlying processor as the receiver, so - /// either one can be used afterwards. The Rust SDK returns `self` here, which a promise - /// cannot express. + /// ```javascript + /// const processor = await new ProcessorAsync(model, licenseKey).withConfig(sampleRate, blockSize) + /// ``` #[napi(ts_return_type = "Promise")] pub fn with_config( &self, @@ -118,9 +118,16 @@ impl ProcessorAsync { }) } - /// Configures the processor for an audio format. Must be called before processing. + /// Configures the processor for the given audio format. + /// + /// Await this method before processing audio. Use {@link Model#getOptimalSampleRate} and + /// {@link Model#getOptimalBlockSize} for the lowest delay. + /// Initialization allocates memory and runs on a libuv worker thread. /// - /// See {@link Processor#initialize}. Allocates, so it runs on a worker. + /// @param sampleRate - Audio sample rate in Hz. + /// @param blockSize - Number of mono samples per block. + /// @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + /// Variable block sizes can add buffering latency. Larger blocks are always rejected. #[napi(ts_return_type = "Promise")] pub fn initialize( &self, @@ -134,40 +141,34 @@ impl ProcessorAsync { }) } - /// Enhances a mono audio block and resolves to the enhanced samples. + /// Enhances a mono audio block and returns a promise for the enhanced samples. /// - /// Unlike {@link Processor#process} this does **not** write into the caller's array. - /// The samples are copied out before the work is queued, so the input stays valid and - /// untouched while the promise is pending, and the result arrives as a new array: + /// The input is copied before work is queued and remains unmodified. The promise + /// resolves to a new `Float32Array` containing the enhanced samples. /// - /// ```js - /// let audio = new Float32Array(blockSize) - /// for (;;) audio = await processor.process(audio) - /// ``` + /// The instance must be initialized first. The block must contain exactly `blockSize` + /// samples, or at most `blockSize` if `variableBlockSize` is enabled. + /// Await each call before submitting the next block. /// - /// The block must be exactly `blockSize` samples, or at most `blockSize` if - /// `variableBlockSize` was enabled. - // The buffer parameter is spelled out because TypeScript widens a bare `Float32Array` - // to `Float32Array`, which does not assign back to a - // `let audio = new Float32Array(n)` and so breaks the reuse loop above. The buffer - // handed to V8 is always a plain, non-shared ArrayBuffer, so the narrower type holds. + /// ```javascript + /// const enhanced = await processor.process(block) + /// ``` + // Specify ArrayBuffer so the result is assignable to a variable inferred from + // `new Float32Array(n)`. The returned buffer is never a SharedArrayBuffer. #[napi(ts_return_type = "Promise>")] pub fn process(&self, audio: Float32Array) -> AsyncTask { AsyncTask::new(ProcessorProcessTask { slot: self.slot.clone(), - // Copied on the JS thread so the worker owns its samples and JS cannot mutate - // them mid-process. A block is a couple of kilobytes, negligible next to running - // the model over it. + // Copy on the JavaScript thread so the worker owns its input. audio: audio.to_vec(), }) } - /// Creates a handle for reading and writing this processor's parameters and state. + /// Returns a promise for a {@link ProcessorContext} to control this processor. /// - /// Asynchronous because it takes the processor lock, which a queued `process` may - /// briefly hold; awaiting keeps that wait off the event loop. The returned handle is - /// the same {@link ProcessorContext} the synchronous class hands out, with the same - /// synchronous methods. + /// Context creation runs on a worker thread because it may wait for processing to + /// release the instance lock. The returned context's methods are synchronous and can + /// be called while audio is being processed. #[napi(ts_return_type = "Promise")] pub fn get_context(&self) -> AsyncTask { AsyncTask::new(ProcessorContextTask { @@ -175,9 +176,14 @@ impl ProcessorAsync { }) } - /// Ends this processor's telemetry session, after which it can no longer process audio. + /// Terminates the telemetry session associated with this processor. + /// + /// Once termination is handled, the processor can no longer process audio. + /// The session also ends when the native object is destroyed. Use this method when + /// termination must be requested at a specific lifecycle event. /// - /// May block, so it runs on a worker. + /// Termination runs on a libuv worker thread because it may block. + /// If another session is still active, termination can complete asynchronously. #[napi(ts_return_type = "Promise")] pub fn terminate_session(&self) -> AsyncTask { AsyncTask::new(ProcessorTerminateTask { @@ -200,9 +206,8 @@ impl Task for ProcessorWithConfigTask { } fn resolve(&mut self, _env: Env, _: ()) -> Result { - // A second JS handle onto the same native processor. The footprint was reported once - // at construction, so nothing is reported here; the last handle's finalizer withdraws - // it. If the processor was disposed mid-flight, this handle starts out disposed too. + // The returned handle shares the native instance and its existing memory report. + // If disposal occurred during initialization, this handle is also disposed. Ok(ProcessorAsync { slot: self.slot.clone(), }) @@ -237,8 +242,7 @@ impl Task for ProcessorProcessTask { type JsValue = Float32Array; fn compute(&mut self) -> Result> { - // Moved out so `resolve` can hand the buffer to V8 without another copy. The task - // runs once, so leaving an empty Vec behind is fine. + // Transfer the task buffer to `resolve` without another allocation. let mut audio = std::mem::take(&mut self.audio); map_err(lock(&self.slot).get_mut()?.process(&mut audio))?; @@ -246,8 +250,7 @@ impl Task for ProcessorProcessTask { } fn resolve(&mut self, _env: Env, audio: Vec) -> Result { - // Hands the allocation to V8 as an external ArrayBuffer, so the enhanced samples - // are not copied again on the way out. + // Transfer the allocation to V8 as an external ArrayBuffer without copying. Ok(Float32Array::new(audio)) } } diff --git a/src/vad.rs b/src/vad.rs index 1e59e23..5f355f9 100644 --- a/src/vad.rs +++ b/src/vad.rs @@ -10,28 +10,32 @@ use crate::{ use napi::{Env, bindgen_prelude::Float32Array, bindgen_prelude::ObjectFinalize}; use napi_derive::napi; -/// Voice activity detection parameters, all changeable while audio is being processed. +/// Configurable voice activity detection parameters. Values can be changed during processing. #[napi] pub enum VadParameter { - /// How long the VAD keeps reporting speech after speech stops, which stabilizes - /// detected -> not-detected transitions. + /// Controls how long the VAD continues reporting speech after speech stops. /// - /// Speech is reported when at least half the blocks in the last - /// `speechHoldDuration * 2` seconds contained speech, so ongoing speech extends the - /// hold. Rounded to the model's window length, so reads may differ from writes. + /// Speech is reported if at least half the blocks processed in the last + /// `speechHoldDuration * 2` seconds contained speech. Additional speech during this + /// period extends the detection period. /// - /// Range 0.0 to 300x the model window length, in seconds. Model-specific default. + /// The duration is rounded to the nearest model window length, so the value read back + /// may differ from the value set. + /// + /// Range: 0.0 to 300 times the model window length, in seconds. Default: model-specific. SpeechHoldDuration = 0, - /// Probability threshold above which a block counts as speech, stabilizing how - /// readily speech is detected at all. + /// Sets the probability threshold for detecting speech in an audio block. + /// + /// A model probability above this threshold counts as speech. /// - /// Range 0.0 - 1.0. Model-specific default. + /// Range: 0.0 to 1.0. Default: model-specific. Sensitivity = 1, - /// How long speech must be present before the VAD reports it, which stabilizes - /// not-detected -> detected transitions. + /// Controls how long speech must be present before the VAD reports speech. + /// + /// The duration is rounded to the nearest model window length, so the value read back + /// may differ from the value set. /// - /// Rounded to the model's window length, so reads may differ from writes. - /// Range 0.0 - 1.0, in seconds. Model-specific default. + /// Range: 0.0 to 1.0 seconds. Default: model-specific. MinimumSpeechDuration = 2, } @@ -45,20 +49,18 @@ impl From for aic_sdk::VadParameter { } } -/// Voice activity detector running a dedicated VAD model. +/// Detects speech using a dedicated VAD model. /// -/// Driven explicitly through {@link Vad#process} and independent of any -/// {@link Processor}; predictions are read through a {@link VadContext}. Enhancement -/// models are rejected. +/// Call {@link Vad#initialize}, then pass mono audio to {@link Vad#process}. +/// Processing leaves the audio unmodified and updates the prediction, which can be read +/// through a {@link VadContext}. /// -/// When enhancement and detection run together, feed this the **original** audio, not the -/// processor's output: enhancement changes the signal the VAD model expects, and stacks -/// the processor's delay onto the prediction. `process` leaves its input untouched, so -/// call it on the same block before `Processor#process`. +/// When using enhancement and detection together, pass the original input to the VAD +/// before calling {@link Processor#process}. Enhanced audio changes the signal seen by +/// the VAD and adds the processor's audio delay to the prediction delay. #[napi(custom_finalize)] pub struct Vad { - // No lock: every method here runs on the JS thread. Only the async class shares its - // slot with tasks on the libuv pool. + // Only the JavaScript thread accesses this slot. Async classes use a shared, locked slot. slot: DisposableSlot>, } @@ -72,10 +74,15 @@ impl ObjectFinalize for Vad { #[napi] impl Vad { - /// Creates a voice activity detector from a dedicated VAD model. + /// Creates a new voice activity detector. /// - /// Telemetry follows the runtime environment; pass `otelConfig` to override it for this - /// instance. + /// Construction is synchronous and throws if creation fails. Call + /// {@link Vad#initialize} before processing audio. + /// + /// @param model - Dedicated VAD model. Other model types are rejected. + /// @param licenseKey - SDK license key from . + /// @param otelConfig - Optional telemetry configuration. When omitted, telemetry follows + /// the runtime environment. #[napi(constructor)] pub fn new( env: Env, @@ -96,19 +103,25 @@ impl Vad { }) } - /// Destroys the native VAD immediately, releasing its memory and telemetry session - /// without waiting for garbage collection. + /// Destroys the native VAD and releases its telemetry session. /// - /// Every later method throws; calling `dispose()` again does nothing. + /// Use this for cleanup at a specific point instead of waiting for garbage collection. + /// After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. #[napi] pub fn dispose(&mut self, env: Env) { self.slot.release(env); } - /// Configures the VAD for an audio format. Must be called before processing. + /// Configures the VAD for the given audio format. + /// + /// Call this method before processing audio. Use {@link Model#getOptimalSampleRate} and + /// {@link Model#getOptimalBlockSize} for the most frequent prediction updates. + /// This method allocates memory; avoid calling it from audio processing callbacks. /// - /// The model's optimal sample rate and block size give the most frequent prediction - /// updates. Allocates, so keep it off the audio path. + /// @param sampleRate - Audio sample rate in Hz. + /// @param blockSize - Number of mono samples per block. + /// @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + /// Variable block sizes can add buffering latency. Larger blocks are always rejected. #[napi] pub fn initialize( &mut self, @@ -123,17 +136,19 @@ impl Vad { ))) } - /// Examines a mono audio block and updates the prediction, leaving the audio unmodified. + /// Processes a mono audio block and updates the VAD prediction without modifying the input. + /// + /// Call {@link Vad#initialize} first. The block must contain exactly `blockSize` samples, + /// or at most `blockSize` if `variableBlockSize` is enabled. #[napi] pub fn process(&mut self, audio: Float32Array) -> Result<()> { - // Read-only, so the safe `Deref` to `&[f32]` covers it. Taking the view by value - // does not copy the caller's samples. + // Borrow the typed-array contents without copying or modifying them. map_err(self.slot.get_mut()?.process(&audio)) } /// Creates a handle for reading predictions and controlling this VAD. /// - /// Each call returns an independent handle onto the same VAD. + /// Each call returns an independent handle to the same VAD. #[napi] pub fn get_context(&self) -> Result { Ok(VadContext { @@ -141,7 +156,14 @@ impl Vad { }) } - /// Ends this VAD's telemetry session, after which it can no longer process audio. + /// Terminates the telemetry session associated with this VAD. + /// + /// Once termination is handled, the VAD can no longer process audio. + /// The session also ends when the native object is destroyed. Use this method when + /// termination must be requested at a specific lifecycle event. + /// + /// This method may block. Avoid calling it from audio processing callbacks. + /// If another session is still active, termination can complete asynchronously. #[napi] pub fn terminate_session(&mut self) -> Result<()> { map_err(self.slot.get_mut()?.terminate_session()) @@ -165,13 +187,13 @@ impl VadContext { map_err(self.inner.set_parameter(parameter.into(), value as f32)) } - /// Reads the current value of a VAD parameter. + /// Returns the current value of a VAD parameter. #[napi] pub fn get_parameter(&self, parameter: VadParameter) -> Result { map_err(self.inner.parameter(parameter.into())).map(f64::from) } - /// Whether speech is currently detected. + /// Returns whether speech is currently detected. /// /// The decision lags its input by {@link VadContext#getPredictionDelay} samples, and /// stops updating if the backing VAD stops being processed. @@ -180,39 +202,49 @@ impl VadContext { self.inner.is_speech_detected() } - /// The model's raw prediction, in the range 0.0 - 1.0. + /// Returns the model's speech probability in the range 0.0 to 1.0. /// - /// Unlike {@link VadContext#isSpeechDetected} this skips the SDK's post-processing - /// (speech hold, sensitivity thresholding), for building your own abstractions on top. - /// The same latency notes apply. + /// This value excludes speech hold, sensitivity thresholding and minimum speech duration. + /// Use it to implement custom detection logic. The prediction delay reported by + /// {@link VadContext#getPredictionDelay} also applies to this value. #[napi] pub fn get_raw_vad_probability(&self) -> f64 { self.inner.raw_vad_probability().into() } - /// How far the prediction lags its input, in samples at the initialized rate. + /// Returns the prediction delay in samples at the configured sample rate. + /// + /// Includes input buffering, STFT and model processing. Non-optimal or variable block + /// sizes can add buffering latency. Convert to milliseconds with + /// `delaySamples * 1000 / sampleRate`. /// - /// Covers input reblocking, STFT and model processing. This delay is **not** applied to - /// the audio (`process` leaves the buffer untouched), so use it to line speech - /// decisions up with the audio timeline. Independent of a processor's audio delay. + /// Use this delay to align speech decisions with the input audio. It is independent of + /// a processor's audio delay; VAD processing does not delay or modify the audio. #[napi] pub fn get_prediction_delay(&self) -> u32 { self.inner.prediction_delay() as u32 } - /// Clears internal state, including the published prediction. + /// Clears internal state and buffers, including the published speech decision and probability. /// - /// Call this on a stream discontinuity or when seeking, to keep earlier audio from - /// causing mispredictions. + /// Call this when the stream is interrupted or when seeking to prevent predictions + /// from using previous audio. The VAD remains initialized with its configured settings. #[napi] pub fn reset(&self) -> Result<()> { map_err(self.inner.reset()) } - /// Swaps in a renewed JWT without interrupting processing. + /// Replaces the bearer token on the running VAD. + /// + /// Use this to refresh a JWT without recreating the instance. Both the original license + /// key and the new token must be JWTs. If this call fails, the previous token remains active. + /// + /// A successful call validates the token's format and applies it immediately. Backend + /// acceptance is checked later. If the backend rejects the token, the SDK retries with + /// backoff; processing is eventually disabled if no accepted token arrives in time. + /// Supply a valid token to recover the session. /// - /// Only works when both the original key and the new token are JWTs. On failure the - /// call is a no-op and the previous token stays active. + /// This method allocates memory and takes a mutex. Avoid calling it from audio processing callbacks. #[napi] pub fn update_bearer_token(&self, token: String) -> Result<()> { map_err(self.inner.update_bearer_token(&token)) diff --git a/src/vad_async.rs b/src/vad_async.rs index e301c41..3bdedcc 100644 --- a/src/vad_async.rs +++ b/src/vad_async.rs @@ -15,27 +15,23 @@ use napi::{ use napi_derive::napi; use std::sync::{Arc, Mutex}; -/// Voice activity detector that keeps its work off the main thread. +/// Voice activity detection for use in async applications. /// -/// The same detection as {@link Vad}, but each call returns a promise and runs on Node's -/// libuv thread pool, so the event loop stays responsive. Predictions are read through a -/// {@link VadContext}, whose methods are all synchronous. +/// Initialization, processing and context creation run on Node's libuv thread pool and +/// return promises. Construction and disposal are synchronous. /// -/// Mirrors `VadAsync` in the Rust SDK. +/// Read predictions through a {@link VadContext}. Pass the original input audio to the +/// VAD before enhancement. /// -/// As with {@link Vad}, feed this the **original** audio when enhancement and detection -/// run together, not a processor's output. +/// ### Threading /// -/// ### Concurrency +/// Use one instance per stream and await each operation before submitting the next. +/// Concurrent calls on one instance are not guaranteed to execute in submission order. +/// Use separate instances to process multiple streams concurrently. /// -/// One instance handles one stream. Do not start a second {@link VadAsync#process} before -/// the first resolves: libuv completes work items out of order, which would desync the -/// stream and scramble the prediction. To watch several streams at once, create several -/// instances. -/// -/// The libuv pool is four threads by default and is shared with `fs`, `dns` and `crypto`. -/// Raise `UV_THREADPOOL_SIZE` before Node starts to run more streams in parallel. -/// `AIC_NUM_THREADS` has no effect: it sizes a rayon pool this binding does not use. +/// The libuv pool defaults to four threads and is shared with filesystem, DNS and crypto +/// work. Set `UV_THREADPOOL_SIZE` before starting Node to change its size. +/// `AIC_NUM_THREADS` does not apply to these bindings. #[napi(custom_finalize)] pub struct VadAsync { slot: Arc>>>, @@ -43,9 +39,9 @@ pub struct VadAsync { impl ObjectFinalize for VadAsync { fn finalize(self, env: Env) -> Result<()> { - // Only the last handle destroys the native object; while other handles or in-flight - // tasks hold an `Arc`, this leaves the object and its footprint report to them. - // A no-op if `dispose()` already ran. + // Release the object and V8 memory estimate only for the last shared handle. + // A pending task can retain the slot beyond finalization; see `DisposableSlot`. + // `release` has no effect if the object was already disposed. if Arc::strong_count(&self.slot) == 1 { lock(&self.slot).release(env); } @@ -55,13 +51,15 @@ impl ObjectFinalize for VadAsync { #[napi] impl VadAsync { - /// Creates a voice activity detector from a dedicated VAD model. + /// Creates a new async voice activity detector. /// - /// Construction is synchronous and throws on failure, as in the Rust SDK; only the - /// audio work runs on a worker thread. + /// Construction is synchronous and throws if creation fails. Await + /// {@link VadAsync#initialize} or {@link VadAsync#withConfig} before processing audio. /// - /// Telemetry follows the runtime environment; pass `otelConfig` to override it for this - /// instance. + /// @param model - Dedicated VAD model. Other model types are rejected. + /// @param licenseKey - SDK license key from . + /// @param otelConfig - Optional telemetry configuration. When omitted, telemetry follows + /// the runtime environment. #[napi(constructor)] pub fn new( env: Env, @@ -87,26 +85,26 @@ impl VadAsync { }) } - /// Destroys the native VAD immediately, releasing its memory and telemetry session - /// without waiting for garbage collection. + /// Destroys the native VAD and releases its telemetry session. + /// + /// Use this for cleanup at a specific point instead of waiting for garbage collection. + /// After disposal, all methods except `dispose()` fail. Repeated disposal has no effect. /// - /// Every later method throws; calling `dispose()` again does nothing. Blocks until - /// in-flight work on the libuv pool finishes. + /// This call blocks the calling thread while a worker holds the instance lock. + /// Queued work that acquires the lock after disposal rejects its promise. #[napi] pub fn dispose(&self, env: Env) { lock(&self.slot).release(env); } - /// Initializes the VAD and resolves to a handle onto it, for chaining off the - /// constructor: + /// Initializes the VAD and returns a promise for a handle to the initialized instance. /// - /// ```js - /// const vad = await new VadAsync(model, licenseKey).withConfig(16000, 160) - /// ``` + /// Uses the same configuration as {@link VadAsync#initialize}. The returned handle and + /// this object share the same native instance; disposing either invalidates both. /// - /// The handle it resolves to drives the same underlying VAD as the receiver, so either - /// one can be used afterwards. The Rust SDK returns `self` here, which a promise cannot - /// express. + /// ```javascript + /// const vad = await new VadAsync(model, licenseKey).withConfig(sampleRate, blockSize) + /// ``` #[napi(ts_return_type = "Promise")] pub fn with_config( &self, @@ -120,9 +118,16 @@ impl VadAsync { }) } - /// Configures the VAD for an audio format. Must be called before processing. + /// Configures the VAD for the given audio format. /// - /// See {@link Vad#initialize}. Allocates, so it runs on a worker. + /// Await this method before processing audio. Use {@link Model#getOptimalSampleRate} and + /// {@link Model#getOptimalBlockSize} for the most frequent prediction updates. + /// Initialization allocates memory and runs on a libuv worker thread. + /// + /// @param sampleRate - Audio sample rate in Hz. + /// @param blockSize - Number of mono samples per block. + /// @param variableBlockSize - Allow blocks shorter than `blockSize`. Defaults to `false`. + /// Variable block sizes can add buffering latency. Larger blocks are always rejected. #[napi(ts_return_type = "Promise")] pub fn initialize( &self, @@ -136,39 +141,33 @@ impl VadAsync { }) } - /// Examines a mono audio block, updates the prediction, and resolves to the same - /// samples unmodified. + /// Updates the VAD prediction and returns a promise for the original mono audio samples. + /// + /// The input is copied before work is queued and remains unmodified. The promise + /// resolves to a new `Float32Array` containing the original samples. /// - /// The samples are copied out before the work is queued, so the caller's array stays - /// valid and untouched while the promise is pending. The block is handed back, instead - /// of the promise resolving to nothing, to match the Rust SDK and to keep a streaming - /// loop reading the same either side of the boundary: + /// The instance must be initialized first. The block must contain exactly `blockSize` + /// samples, or at most `blockSize` if `variableBlockSize` is enabled. + /// Await each call before submitting the next block. /// - /// ```js - /// let audio = new Float32Array(blockSize) - /// for (;;) { - /// audio = await vad.process(audio) - /// console.log(context.isSpeechDetected()) - /// } + /// ```javascript + /// const audio = await vad.process(block) /// ``` // See the note on {@link ProcessorAsync#process} for why the buffer type is spelled out. #[napi(ts_return_type = "Promise>")] pub fn process(&self, audio: Float32Array) -> AsyncTask { AsyncTask::new(VadProcessTask { slot: self.slot.clone(), - // Copied on the JS thread so the worker owns its samples and JS cannot mutate - // them mid-process. A block is a couple of kilobytes, negligible next to running - // the model over it. + // Copy on the JavaScript thread so the worker owns its input. audio: audio.to_vec(), }) } - /// Creates a handle for reading predictions and controlling this VAD. + /// Returns a promise for a {@link VadContext} to control this VAD and read predictions. /// - /// Asynchronous because it takes the VAD lock, which a queued `process` may briefly - /// hold; awaiting keeps that wait off the event loop. The returned handle is the same - /// {@link VadContext} the synchronous class hands out, whose methods are synchronous, - /// so a prediction can be read from inside an audio callback. + /// Context creation runs on a worker thread because it may wait for processing to + /// release the instance lock. The returned context's methods are synchronous and can + /// be called while audio is being processed. #[napi(ts_return_type = "Promise")] pub fn get_context(&self) -> AsyncTask { AsyncTask::new(VadContextTask { @@ -176,9 +175,14 @@ impl VadAsync { }) } - /// Ends this VAD's telemetry session, after which it can no longer process audio. + /// Terminates the telemetry session associated with this VAD. + /// + /// Once termination is handled, the VAD can no longer process audio. + /// The session also ends when the native object is destroyed. Use this method when + /// termination must be requested at a specific lifecycle event. /// - /// May block, so it runs on a worker. + /// Termination runs on a libuv worker thread because it may block. + /// If another session is still active, termination can complete asynchronously. #[napi(ts_return_type = "Promise")] pub fn terminate_session(&self) -> AsyncTask { AsyncTask::new(VadTerminateTask { @@ -201,9 +205,8 @@ impl Task for VadWithConfigTask { } fn resolve(&mut self, _env: Env, _: ()) -> Result { - // A second JS handle onto the same native VAD. The footprint was reported once at - // construction, so nothing is reported here; the last handle's finalizer withdraws - // it. If the VAD was disposed mid-flight, this handle starts out disposed too. + // The returned handle shares the native instance and its existing memory report. + // If disposal occurred during initialization, this handle is also disposed. Ok(VadAsync { slot: self.slot.clone(), }) @@ -238,8 +241,7 @@ impl Task for VadProcessTask { type JsValue = Float32Array; fn compute(&mut self) -> Result> { - // Moved out so `resolve` can hand the buffer to V8 without another copy. The task - // runs once, so leaving an empty Vec behind is fine. + // Transfer the task buffer to `resolve` without another allocation. let audio = std::mem::take(&mut self.audio); map_err(lock(&self.slot).get_mut()?.process(&audio))?; @@ -247,8 +249,7 @@ impl Task for VadProcessTask { } fn resolve(&mut self, _env: Env, audio: Vec) -> Result { - // Hands the allocation to V8 as an external ArrayBuffer, so the block is not copied - // again on the way out. + // Transfer the allocation to V8 as an external ArrayBuffer without copying. Ok(Float32Array::new(audio)) } }