Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
139 changes: 72 additions & 67 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ async function main() {
const modelPath = await Model.download('quail-vf-2.2-s-16khz', './models')
const model = Model.fromFile(modelPath)

// The model's own settings give the lowest delay
// Use the model's optimal configuration for the lowest delay
const sampleRate = model.getOptimalSampleRate()
const blockSize = model.getOptimalBlockSize(sampleRate)

Expand All @@ -45,30 +45,32 @@ main()
```

Processing is mono. For multichannel audio, mix down to mono or use one processor per
channel.
channel. Initialize each instance before passing audio. Blocks must contain exactly
`blockSize` samples. Set the third `initialize` argument, `variableBlockSize`, to `true`
to allow shorter blocks; longer blocks are always rejected. This also applies to VAD
processing and analyzer buffering.

The snippets below are fragments, and each one leaves out that surrounding `async function`
for brevity. Keep it: `await` at the top level of a CommonJS file is a syntax error, and
several calls here are asynchronous.
The following snippets assume an enclosing `async` function. CommonJS files do not support
top-level `await`.

Runnable scripts for enhancement, VAD, analysis and whole-file processing, synchronous and
async, are in [`examples/`](examples).

## Models

Available models and their ids are listed at
[artifacts.ai-coustics.io](https://artifacts.ai-coustics.io). Each class accepts exactly one
family of models and throws for the rest:
Available models and their IDs are listed at
[artifacts.ai-coustics.io](https://artifacts.ai-coustics.io). Each class accepts the model types
listed below and rejects other types:

| Class | Accepted models |
| ----------------------------- | ------------------- |
| `Processor`, `ProcessorAsync` | enhancement, bypass |
| `Vad`, `VadAsync` | dedicated VAD |
| `Analyzer` | analysis |

`Model.download` resolves to the model's path and runs off the event loop. Model files are
memory-mapped rather than read into memory, so keep the file in place while anything created
from it is alive.
`Model.download` runs on Node's libuv thread pool and returns a promise for the model's path.
`Model.fromFile` memory-maps the file. Do not modify or delete it while the model or any
instance created from it is still alive.

## Enhancement

Expand All @@ -84,40 +86,38 @@ context.setParameter(ProcessorParameter.Bypass, 0)

console.log(context.getParameter(ProcessorParameter.EnhancementLevel))

// Samples of delay the processor adds, for lining the output up with other streams
// Get the audio delay in samples to align the output with other streams
console.log(context.getAudioDelay())

// Clear internal state on a stream discontinuity or seek
context.reset()
```

## Off the main thread
## Async processing

`ProcessorAsync` and `VadAsync` do the same work as `Processor` and `Vad`, but each call
returns a promise and runs on a worker thread, so the event loop stays responsive. Use them
when other work shares that loop, as in a server handling live streams alongside its sockets
and HTTP. The synchronous classes are the pick on a dedicated audio thread or in a batch
script,
where nothing else needs the loop and a promise per block would only add overhead.
`ProcessorAsync` and `VadAsync` run initialization and processing on Node's libuv thread pool.
Await these calls to keep the event loop available for other work, such as network requests.
Construction and disposal remain synchronous.

Use `Processor` or `Vad` on a dedicated worker thread, or in a batch script that can block
the calling thread. These classes process each block without a promise or input copy.

```javascript
const { Model, ProcessorAsync } = require('@ai-coustics/aic-sdk')

const processor = await new ProcessorAsync(model, process.env.AIC_SDK_LICENSE).withConfig(sampleRate, blockSize)

// Unlike the synchronous class, this does not write into the caller's array. It resolves
// to the enhanced samples, so a streaming loop reuses the same variable
let audio = new Float32Array(blockSize)
for (;;) {
audio = await processor.process(audio)
}
// The input remains unmodified; the promise resolves to the enhanced samples.
const audio = new Float32Array(blockSize)
const enhanced = await processor.process(audio)
```

The input is copied before the work is queued, so it stays valid and untouched while the
promise is pending. `VadAsync.process` hands the block back unmodified in the same way.
Both async classes copy the input before queuing work and return a new `Float32Array`.
`ProcessorAsync.process` returns enhanced samples; `VadAsync.process` returns the original
samples. The caller's input remains unmodified.

Context handles are awaited but their methods stay synchronous, so a prediction or parameter
can still be read from inside an audio callback:
Context creation is asynchronous. The returned context's methods are synchronous and can
be used while processing runs:

```javascript
const context = await processor.getContext()
Expand All @@ -126,9 +126,9 @@ context.setParameter(ProcessorParameter.EnhancementLevel, 0.8)

### Running several streams

One instance handles one stream. Do not start a second `process` on the same instance before
the first resolves: worker threads complete out of order, which would desync the stream. Give
each stream its own instance instead.
Use one instance per stream and await each operation before submitting the next. Concurrent
calls on one instance are not guaranteed to execute in submission order. Separate instances
can process streams concurrently.

```javascript
const processors = await Promise.all(
Expand All @@ -146,9 +146,7 @@ Node starts:
UV_THREADPOOL_SIZE=16 node server.js
```

The core SDK's own `AIC_NUM_THREADS` has no effect here: it sizes a thread pool these
bindings deliberately do not use, so that all audio work stays on the pool Node already
manages.
`AIC_NUM_THREADS` does not apply to these bindings; async audio work uses libuv.

## Voice activity detection

Expand Down Expand Up @@ -182,7 +180,8 @@ When enhancement and detection run together, feed the VAD the **original** input
processor's output. Enhancement is designed to change the signal, so detecting on its output
means running the VAD on audio that no longer matches what its model expects, and it stacks
the processor's delay on top of the prediction delay. Because `vad.process` leaves its input
untouched, running both on the same block is enough:
untouched, run it before enhancement. Configure both instances with the same sample rate and
block size:

```javascript
vad.process(block) // reads the block
Expand All @@ -196,13 +195,13 @@ context.getAudioDelay() // enhanced audio lags the input by this many samples
vadContext.getPredictionDelay() // the VAD decision lags the same input by this many
```

The prediction delay is not applied to the audio; use it to line speech decisions up with
the audio timeline.
The prediction delay is not applied to the audio. Use it to align speech decisions with
the input audio.

## Analysis

Analysis models score audio quality. Buffering is cheap enough for the audio path;
`analyze` runs the model and is not.
Analysis models score audio quality. Collect audio with `buffer`, then run the model with
`analyzeAsync` or `analyze`.

```javascript
const { Model, Analyzer } = require('@ai-coustics/aic-sdk')
Expand All @@ -220,19 +219,21 @@ const result = await analyzer.analyzeAsync()
console.log(result.riskScore, result.noise, result.speakerReverb)
```

`buffer` is cheap enough for the audio path; running the model is not, so it is a separate
call with two forms. `analyzeAsync` is the one to reach for in a server: analysis is
occasional, so the promise costs nothing next to the model, and the event loop stays free.
`analyze` blocks and suits a CLI or a worker thread.
`buffer` is synchronous and does not acquire the analyzer lock. Audio collection can
continue while `analyzeAsync` runs on a worker thread. Use `analyze` when blocking the
calling thread is acceptable, such as in a CLI or dedicated worker.

Only that one call moves off-thread, which is why there is no `AnalyzerAsync` class to match
`ProcessorAsync`. `buffer` stays synchronous, takes no lock, and can be called while an
analysis is still running. The SDK guarantees the collector and analyzer halves are safe to
use concurrently. The analyzer's other methods do wait for a pending analysis to finish.
Analysis uses a fixed duration of audio determined by the model. Older samples are discarded
as new audio arrives. If insufficient audio has been collected, analysis pads the remaining
input with silence.

Every score runs 0.0 - 1.0. Except `speakerLoudness`, lower means less problematic audio.
`riskScore` is the headline number: how likely this audio is to break downstream models such
as speech-to-text, VAD or turn-taking.
Calls to `analyze`, `reset`, `updateBearerToken`, `terminateSession` and `dispose` acquire the
analyzer lock and may block while async analysis is running. `initialize` only configures
the collector.

All scores range from 0.0 to 1.0. For every field except `speakerLoudness`, lower values
indicate less problematic audio. `riskScore` predicts the likelihood of failure in downstream
models such as speech-to-text, VAD or turn-taking.

## Telemetry

Expand All @@ -247,10 +248,11 @@ const processor = new Processor(model, licenseKey, {
})
```

A session is closed when its object is garbage collected. Because GC timing is not
guaranteed, every processor, VAD and analyzer also exposes `terminateSession()` for
lifecycle events; afterwards the object can no longer process audio. On `ProcessorAsync`
and `VadAsync` it returns a promise, since it may block.
A telemetry session ends when its native object is destroyed. Use `terminateSession()` to
request termination at a specific lifecycle event. Once termination is handled, processors
and VADs can no longer process audio, and analyzers can no longer analyze buffered audio.
On `ProcessorAsync` and `VadAsync`, termination runs on a libuv worker thread and returns
a promise because it may block.

If your license key is a JWT, refresh it in place instead of rebuilding the object:

Expand All @@ -261,13 +263,12 @@ context.updateBearerToken(renewedJwt)
## Memory management

`Model`, `Processor`, `ProcessorAsync`, `Vad`, `VadAsync` and `Analyzer` hold large native
allocations behind small JavaScript objects. The binding reports each object's native
footprint to V8 (`napi_adjust_external_memory`), so the garbage collector applies the right
amount of pressure and reclaims dropped instances promptly instead of letting native memory
grow unbounded.
allocations behind small JavaScript objects. The binding reports estimated native memory
usage to V8 so the garbage collector can account for these allocations. This influences
collection frequency but does not guarantee when an object will be released.

For deterministic cleanup, every one of these classes also exposes `dispose()`, which
destroys the native object immediately instead of waiting for garbage collection:
Use `dispose()` to release native resources at a specific point instead of waiting for
garbage collection:

```javascript
const processor = new Processor(model, licenseKey)
Expand All @@ -279,10 +280,14 @@ try {
}
```

After `dispose()`, every method on the object throws; calling `dispose()` again does
nothing. On the async classes it blocks until in-flight work on the libuv pool finishes.
After `dispose()`, all methods except `dispose()` fail; repeated disposal has no effect.
Async methods reject their promises. Disposal blocks the calling thread if a worker holds
the native object's lock. Queued work that acquires the lock after disposal fails.

Disposing a model releases its reference to the model data. Processors, VADs and analyzers
created from it retain their own references and remain usable.

Two things to know about cleanup timing:
Cleanup timing also affects native memory usage:

- Native cleanup runs on the event loop when the object is finalized, not synchronously at
garbage collection. Finalizers run on event-loop turns, so batches that create many of
Expand All @@ -291,9 +296,9 @@ Two things to know about cleanup timing:
`setTimeout` or `setImmediate` are. A tight synchronous loop that creates thousands of
objects accumulates their native memory for the duration of the loop; create these
objects per unit of work behind real async boundaries, or reuse a single instance.
- RSS reflects the peak of simultaneously live (or not-yet-finalized) instances: the
allocator reuses freed native memory rather than returning it to the OS, so a burst of N
concurrent instances costs about N x their footprint even after they are dropped.
- RSS may remain elevated after disposal because the allocator can retain freed memory
for reuse. Peak memory use depends on the number of concurrently live instances,
including objects awaiting finalization.

## Development

Expand Down
4 changes: 2 additions & 2 deletions __test__/common.ts
Original file line number Diff line number Diff line change
Expand Up @@ -27,9 +27,9 @@ export function modelPath(kind: ModelKind): string {
}

/**
* The license key the SDK needs to construct anything.
* Returns the SDK license key used by the tests.
*
* Throws up front, so a missing key is not reported later as a license error.
* Throws if AIC_SDK_LICENSE is missing.
*/
export function licenseKey(): string {
const key = process.env.AIC_SDK_LICENSE
Expand Down
7 changes: 3 additions & 4 deletions __test__/dispose.spec.ts
Original file line number Diff line number Diff line change
Expand Up @@ -190,8 +190,7 @@ test('analyzer dispose waits for an in-flight analyzeAsync to finish', async (t)
analyzer.dispose()
const blockedMs = elapsedMs(disposeStarted)

// dispose() waited for the lock instead of pulling the analyzer out from under the
// worker, so the analysis produced a real result.
// The worker completed analysis before disposal acquired the lock.
const result = await inFlight
t.is(typeof result.riskScore, 'number', 'the in-flight analysis must complete, not be cancelled')

Expand All @@ -210,8 +209,8 @@ test('a withConfig handle outlives its original handle being collected', async (
const sampleRate = model.getOptimalSampleRate()
const blockSize = model.getOptimalBlockSize(sampleRate)

// The chaining form leaves the constructor's handle as garbage. Its finalizer must
// withdraw only its own footprint report, not destroy the shared native processor.
// The constructor's temporary handle may be finalized after `withConfig` resolves.
// Its finalizer must preserve the shared processor and its single memory report.
const processor = await new ProcessorAsync(model, licenseKey()).withConfig(sampleRate, blockSize)

// Push V8 towards a GC so the dropped handle's finalizer runs.
Expand Down
18 changes: 7 additions & 11 deletions __test__/index.spec.ts
Original file line number Diff line number Diff line change
Expand Up @@ -44,8 +44,8 @@ test('sdk expects the model version the fixtures are published under', (t) => {
test('model exposes its id and optimal settings', (t) => {
const model = enhancementModel()

// The model file's own id carries the build hash and version, e.g.
// `quail-vf-2.2-s-16khz-gf70x7zf-v14`, so it extends the manifest id used to download it.
// The model file's own ID carries the build hash and version, e.g.
// `quail-vf-2.2-s-16khz-gf70x7zf-v14`, so it extends the manifest ID used to download it.
t.true(
model.getId().startsWith(TEST_MODELS.enhancement.id),
`expected ${model.getId()} to start with ${TEST_MODELS.enhancement.id}`,
Expand Down Expand Up @@ -175,9 +175,8 @@ test('async processor matches the sync processor sample for sample', async (t) =
model.getOptimalBlockSize(model.getOptimalSampleRate()),
)

// Same model, settings and input, and both start from a fresh state, so moving the work
// to a worker thread and copying the block across must not change the output. Four
// blocks, so a drift in the processor's internal state would show up too.
// Compare four blocks from fresh processors with identical input and settings.
// Async scheduling and buffer copies must preserve output and state transitions.
for (let block = 0; block < 4; block += 1) {
const audio = ramp(blockSize)
const expected = audio.slice()
Expand Down Expand Up @@ -216,8 +215,7 @@ test('async processors on separate instances can run concurrently', async (t) =>
const sampleRate = model.getOptimalSampleRate()
const blockSize = model.getOptimalBlockSize(sampleRate)

// Parallelism is across instances, never within one: libuv completes work items out of
// order, so overlapping calls on a single instance would desync the stream.
// Use independent processors for concurrency; calls within each stream must be ordered.
const processors = await Promise.all(
Array.from({ length: 4 }, () => new ProcessorAsync(model, licenseKey()).withConfig(sampleRate, blockSize)),
)
Expand Down Expand Up @@ -408,10 +406,8 @@ test('audio can be buffered while an analysis is in flight', async (t) => {
t.is(typeof result.riskScore, 'number')
})

// 'async processor rejects instead of throwing' covers async errors arriving as a
// rejection, because processing before initialize is a certain error. The analyzer has no
// equally certain one: nobody has checked whether analyzing before initialize errors or
// returns silence-padded scores, so this file asserts nothing about it.
// Async rejection behavior is covered by the processor test, which processes before
// initialization. No equivalent analyzer failure is assumed here.

test('each class rejects the wrong model type', (t) => {
const key = licenseKey()
Expand Down
2 changes: 1 addition & 1 deletion __test__/models.ts
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ export interface TestModel {
* Model fixtures the end-to-end tests run against.
*
* `scripts/fetch-test-models.mjs` resolves these through `Model.download()`, which
* re-fetches the manifest and pulls the newest compatible model version. The model file
* fetches the manifest and selects the latest compatible model version. The model file
* format version is tied to the SDK version, so `getCompatibleModelVersion()` reports
* the version the built addon expects, and the `sdk expects the model version the
* fixtures are published under` test asserts the fixtures still match it.
Expand Down
Loading