Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion .claude/skills/audio-nodes/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,7 +109,7 @@ Settle algorithm (allocation-free):
Key invariants:
- **Pull from `processableState_`, never `AudioNode::isProcessable()`.** A tail-bearing node (Delay/Convolver/Biquad) overrides `isProcessable()` to stay `true` while its tail drains after a disconnect; using that for the pull would wrongly re-activate its whole upstream cone. The tail node stays scheduled via that override; its `processableState_` is `NOT_PROCESSABLE`, so it correctly does not pull upstream.
- **`disable()` is sticky.** `AudioNode::disable()` sets `NOT_PROCESSABLE` **and** `alwaysNotProcessable_ = true`, so a finished source still wired to a live consumer is not re-activated by the every-quantum pull. Sources call `disable()` from the audio thread when playback finishes.
- **DelayReader → DelayWriter** have no audio edge (they share a ring buffer). `Graph::linkNodes(reader, writer)` records a processable-link, mirrored onto `AudioGraph::Node::link_head`. Settle follows links so pulling the reader also pulls the writer and the writer's inputs. Links are NOT part of the topological sort (that would create a cycle for feedback delays).
- **DelayReader → DelayWriter** have no audio edge (they share a ring buffer). `Graph::linkNodes(reader, writer)` records a processable-link, mirrored onto `AudioGraph::Node::link_head`. Settle follows links so pulling the reader also pulls the writer and the writer's inputs. The toposort also treats a link as an ordering constraint (target before holder) so the writer runs before the reader within a quantum; without that the writer, being an audio sink, would be sorted last and every delay shorter than one quantum would lose samples (the `DelayLine` read snapshot does not make order irrelevant). When the link would close a cycle (feedback delay), the sort drops the constraint for that cycle and orders it by edges only, which is fine because a delay inside a cycle is at least one quantum. Inside such a cycle the reader therefore still runs first; `DelayLine::readerRanThisQuantum()` tells the writer so, and it then writes at least one quantum ahead (the spec's in-cycle clamp). The ring holds `max(maxDelay frames, quantum) + quantum` so neither a delay at `maxDelayTime` nor the in-cycle clamp wraps onto the current read window. Composite host objects must map `getInput()` to the node that receives audio (writer) and `getOutput()` to the node that emits it (reader); `AudioNodeHostObject::connect` uses `source.getOutput()` → `dest.getInput()`.

---

Expand Down Expand Up @@ -353,6 +353,12 @@ All graph mutations are queued via `AudioGraphManager` using its own SPSC channe
### Settable channel attributes (channelCount / channelCountMode / channelInterpretation)
These are mutable after construction. `AudioNode` (core) exposes virtual `setChannelCount` / `setChannelCountMode` / `setChannelInterpretation`. `channelCount` and `channelCountMode` are read only on the host thread during negotiation, so the JSI setter updates the core field directly then calls `HostNode::renegotiate()` → `Graph::renegotiateNodeChannels()` → `HostGraph::renegotiateNodeChannels()` (reuses `collectNegotiations` + an `AGEvent` buffer swap, self-drain aware when there is no audio/render consumer — offline construction/suspend and realtime suspended/stopped windows). When `AudioBufferSourceNode` `setBuffer` changes channel width, update `channelCount_` on the host thread then `renegotiate()` so MAX/CLAMPED_MAX downstream nodes update; the audio event still installs the prebuilt buffer (no audio-thread alloc). `channelInterpretation` is read on the audio thread in `processInputs` (`getInputBuffer()->sum(*input, channelInterpretation_)`), so it MUST be applied via `scheduleAudioEvent`, not mutated directly.

### Output width that differs from the negotiated input width
Negotiation computes an input width per node, asks `getUpstreamChannelCount(negotiated)` (JS thread) what the node presents downstream, and allocates one buffer via `negotiateBufferChannelCount(negotiated)` (JS thread, default identity). Two patterns exist:
- **Fixed output width, input-dependent processing** (`StereoPannerNode`): keep the default buffer width (the input width, needed to pick the mono pan law) and override `getNegotiatedBuffer()`/`setNegotiatedBuffer()` to route negotiation at the input buffer plus `getOutputBuffer()`/`getOutput()` for a separately owned output buffer.
- **Output width computable on the JS thread** (`ConvolverNode`): override `negotiateBufferChannelCount()` to size the single in-place buffer at the output width and `getUpstreamChannelCount()` to report the same. Anything the width depends on that arrives later (Convolver's IR) is handed to the node on the host thread *before* `renegotiate()` (`setImpulseResponseForNegotiation`). Convolver instances carry per-stream state, so a mono IR is installed with one and `negotiateBufferChannelCount` schedules the second lazily when a stereo input first appears; the audio-thread routing clamps to `min(buffer channels, convolvers)` for the window between the buffer event and that append event. Note for tests: an idle `OfflineAudioContext` drains scheduled audio events synchronously, so such lazily scheduled work lands immediately. The IR event and the buffer event travel on different queues, so audio-thread routing must clamp to the buffer's real channel count for the quantum they disagree. The base mixer mixes sources straight into the buffer at the buffer's width, which is wrong when the buffer is wider than the computed input (a quad source with `channelCount: 1` must fold to mono first, with mono gains). Such a node overrides `mixInputs()` (new `AudioNode` hook) to mix at the negotiated width into a preallocated scratch and copy the result to every buffer channel.
Source nodes report their `channelCount_` attribute as output width; `OscillatorNode`/`ConstantSourceNode` pass a mono copy of their options to the base (`withMonoOutput()`) while the host object keeps the spec attribute value.

### Idle-node stale-buffer zeroing (settleProcessableState)
`AudioGraph::iter()` filters to `isProcessable()` nodes, so a node that has gone idle (e.g. a finished source) is skipped and its output buffer is NOT refreshed — it keeps the samples from an earlier quantum. Downstream consumers still read that buffer via `getOutput()` when collecting inputs, which would re-sum ghost echoes every quantum (this broke the `audionode-channel-rules` ~170-node WPT test).

Expand Down
6 changes: 6 additions & 0 deletions packages/audiodocs/docs/effects/convolver-node.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,12 @@ Inherits all properties from [`AudioNode`](../core/audio-node.mdx#properties).
| `buffer` | [`AudioBuffer`](../sources/audio-buffer.mdx) | Associated AudioBuffer. Setting it throws `NotSupportedError` unless the buffer has 1, 2 or 4 channels and the same sample rate as the context. |
| `normalize` | `boolean` | Whether the impulse response from the buffer will be scaled by an equal-power normalization when the buffer attribute is set. |

:::info Channel constraints
The input is mixed down to mono or stereo before convolution. `channelCount` cannot be set above `2` and `channelCountMode` cannot be set to `'max'`; either attempt, in the constructor options or via the setter, throws `NotSupportedError`.

The output is mono only when a mono input meets a mono impulse response. Every other supported combination (a stereo input, or a 2- or 4-channel impulse response) produces stereo. A 4-channel buffer performs "true stereo" convolution: channels 0 and 1 are driven by the left input and channels 2 and 3 by the right input, with even channels summed into the left output and odd channels into the right. Without a buffer the node outputs a single channel of silence.
:::

:::caution
Linear convolution is a heavy computational process, so if your audio has some weird artefacts that should not be there, try to decrease the duration of impulse response buffer.
:::
Expand Down
2 changes: 1 addition & 1 deletion packages/audiodocs/docs/effects/delay-node.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ Inherits all properties from [`AudioNodeOptions`](../core/audio-node.mdx#audiono

| Parameter | Type | Default | |
| :---: | :---: | :----: | :---- |
| `maxDelayTime` <Optional /> | `number` | `1.0` | Maximum amount of time, in seconds, to buffer delayed values. |
| `maxDelayTime` <Optional /> | `number` | `1.0` | Maximum amount of time, in seconds, to buffer delayed values. Must be greater than `0` and less than `180`, otherwise the constructor throws `NotSupportedError`. |
| `delayTime` <Optional /> | `number` | `0.0` | Initial value for [`delayTime`](./delay-node.mdx#properties). |

You can also create a `DelayNode` via the [`BaseAudioContext.createDelay(maxDelayTime?: number)`](../core/base-audio-context.mdx#createdelay) factory method, which uses default values when called without arguments.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@

#include <audioapi/utils/ThreadPool.hpp>

#include <algorithm>
#include <memory>
#include <utility>
#include <vector>
Expand Down Expand Up @@ -45,6 +46,8 @@ JSI_PROPERTY_SETTER_IMPL(ConvolverNodeHostObject, normalize) {

JSI_HOST_FUNCTION_IMPL(ConvolverNodeHostObject, setBuffer) {
if (!args[0].isObject()) {
clearBuffer();
thisValue.asObject(runtime).setExternalMemoryPressure(runtime, getMemoryPressure());
return jsi::Value::undefined();
}

Expand All @@ -57,8 +60,20 @@ JSI_HOST_FUNCTION_IMPL(ConvolverNodeHostObject, setBuffer) {
return jsi::Value::undefined();
}

void ConvolverNodeHostObject::clearBuffer() {
irBytes_ = 0;

convolverNode_->scheduleAudioEvent(
[node = convolverNode_](BaseAudioContext & /*context*/) { node->clearBuffer(); });

// The output collapses to one channel of silence; downstream widths follow.
convolverNode_->setImpulseResponseForNegotiation(nullptr, 0);
renegotiate();
}

void ConvolverNodeHostObject::setBuffer(const std::shared_ptr<AudioBuffer> &buffer) {
if (buffer == nullptr) {
clearBuffer();
return;
}

Expand All @@ -72,23 +87,19 @@ void ConvolverNodeHostObject::setBuffer(const std::shared_ptr<AudioBuffer> &buff
}

auto threadPool = std::make_shared<ConvolverThreadPool>(4);
const size_t irChannels = copiedBuffer->getNumberOfChannels();
std::vector<std::unique_ptr<Convolver>> convolvers;
for (size_t i = 0; i < copiedBuffer->getNumberOfChannels(); ++i) {
AudioArray channelData(*copiedBuffer->getChannel(i));
convolvers.push_back(std::make_unique<Convolver>());
convolvers.back()->init(RENDER_QUANTUM_SIZE, channelData, copiedBuffer->getSize());
}
if (copiedBuffer->getNumberOfChannels() == 1) {
// add one more convolver, because right now input is always stereo
AudioArray channelData(*copiedBuffer->getChannel(0));
convolvers.push_back(std::make_unique<Convolver>());
convolvers.back()->init(RENDER_QUANTUM_SIZE, channelData, copiedBuffer->getSize());
convolvers.reserve(irChannels);
for (size_t channel = 0; channel < irChannels; ++channel) {
convolvers.push_back(ConvolverNode::makeConvolver(*copiedBuffer, channel));
}

constexpr size_t kMaxOutputChannels = 2;
auto internalBuffer = std::make_shared<DSPAudioBuffer>(
RENDER_QUANTUM_SIZE * 2, convolverNode_->getChannelCount(), copiedBuffer->getSampleRate());
RENDER_QUANTUM_SIZE * 2, kMaxOutputChannels, copiedBuffer->getSampleRate());
// Room for the second convolver a mono response may gain later
auto intermediateBuffer = std::make_shared<DSPAudioBuffer>(
RENDER_QUANTUM_SIZE, convolvers.size(), copiedBuffer->getSampleRate());
RENDER_QUANTUM_SIZE, std::max(irChannels, kMaxOutputChannels), copiedBuffer->getSampleRate());

struct SetupData {
std::shared_ptr<AudioBuffer> buffer;
Expand Down Expand Up @@ -118,5 +129,10 @@ void ConvolverNodeHostObject::setBuffer(const std::shared_ptr<AudioBuffer> &buff
context.getDisposer()->dispose(std::move(setupData));
};
convolverNode_->scheduleAudioEvent(std::move(event));

// Negotiation on this thread must see the new IR before the graph recomputes
// what this node presents downstream.
convolverNode_->setImpulseResponseForNegotiation(copiedBuffer, convolvers.size());
renegotiate();
}
} // namespace audioapi
Original file line number Diff line number Diff line change
Expand Up @@ -33,5 +33,6 @@ class ConvolverNodeHostObject : public AudioNodeHostObject {
bool normalize_;
size_t irBytes_ = 0;
void setBuffer(const std::shared_ptr<AudioBuffer> &buffer);
void clearBuffer();
};
} // namespace audioapi
Original file line number Diff line number Diff line change
Expand Up @@ -25,11 +25,6 @@ DelayNodeHostObject::DelayNodeHostObject(
delayTimeParam_ =
std::make_shared<AudioParamHostObject>(graph_, node_, delayNode_->getDelayTimeParam());

auto delayBuffer = std::make_shared<AudioBuffer>(
static_cast<size_t>(options.maxDelayTime * context->getSampleRate() + 1),
channelCount_,
context->getSampleRate());

// order has to be preserved because adding cycle would not change their order in the graph
delayReaderHostNode_ =
std::make_shared<DelayReaderHostNode>(graph_, std::move(delayNode_->delayReader_));
Expand All @@ -47,24 +42,22 @@ DelayNodeHostObject::DelayNodeHostObject(
addGetters(JSI_EXPORT_PROPERTY_GETTER(DelayNodeHostObject, delayTime));
}

std::shared_ptr<utils::graph::HostNode> DelayNodeHostObject::getInput(int /*outputIndex*/) {
return delayReaderHostNode_;
std::shared_ptr<utils::graph::HostNode> DelayNodeHostObject::getInput(int /*inputIndex*/) {
return delayWriterHostNode_;
}

std::shared_ptr<utils::graph::HostNode> DelayNodeHostObject::getOutput(int /*inputIndex*/) {
return delayWriterHostNode_;
std::shared_ptr<utils::graph::HostNode> DelayNodeHostObject::getOutput(int /*outputIndex*/) {
return delayReaderHostNode_;
}

JSI_PROPERTY_GETTER_IMPL(DelayNodeHostObject, delayTime) {
return jsi::Object::createFromHostObject(runtime, delayTimeParam_);
}

size_t DelayNodeHostObject::getMemoryPressure() const {
const float maxDelaySeconds = delayNode_->getDelayTimeParam()->getMaxValue();
const float sampleRate = delayNode_->getContextSampleRate();
// The delay line ring buffer dominates: (maxDelay * sr + 1) frames * channels * float.
// The delay line ring buffer dominates: ring frames * channels * float.
const size_t ringBytes =
static_cast<size_t>(maxDelaySeconds * sampleRate + 1) * channelCount_ * sizeof(float);
delayNode_->delayLine_->getBuffer()->getSize() * channelCount_ * sizeof(float);
// Base `audioBuffer_` from AudioNodeHostObject::getMemoryPressure(), plus
// the reader/writer AudioNode sub-nodes (each owns its own RQ audioBuffer_)
// and the delayTime AudioParam.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -129,6 +129,15 @@ class AudioNode : public utils::graph::GraphObject, public std::enable_shared_fr
return negotiatedChannelCount;
}

/// @brief Channel count of the buffer negotiation allocates for this node,
/// given the negotiated input width. The default returns the input width,
/// which is right for a node that processes in place. A node whose output
/// is wider than its input overrides it to get the buffer at output width
/// and then mixes at the input width itself in `mixInputs`.
[[nodiscard]] virtual size_t negotiateBufferChannelCount(size_t negotiatedChannelCount) {
return negotiatedChannelCount;
}

/// @note JS Thread only
[[nodiscard]] bool requiresTailProcessing() const;

Expand Down Expand Up @@ -252,13 +261,22 @@ class AudioNode : public utils::graph::GraphObject, public std::enable_shared_fr
updateTailStateForQuantum(inputs, numFrames);
}

mixInputs(inputs);

processNode(numFrames);
}

/// @brief Mixes the upstream outputs into the input buffer: zeroes it, then
/// sums every input using this node's channelInterpretation, so the mix
/// happens at the buffer's channel count. A node whose buffer is wider than
/// its negotiated input width overrides this to mix at that width first.
/// @note Audio Thread only
virtual void mixInputs(const std::vector<const DSPAudioBuffer *> &inputs) {
getInputBuffer()->zero();

for (const DSPAudioBuffer *input : inputs) {
getInputBuffer()->sum(*input, channelInterpretation_);
}

processNode(numFrames);
}

/// @brief Returns the tail length in audio frames for the current node
Expand Down
Loading
Loading