
SoundA2Z
Private repoA control plane for the audio you already own.
Per-app volume, sample-accurate multi-room zones, and one API an assistant can actually drive. Not a media player.
- crates
- 9
- lines of Rust
- 112k
- tests
- 689
- operating systems
- 3
crates
lines of Rust
tests
operating systems
Three problems, and no single tool solving them
Per-app volume is solved on exactly one operating system. Windows has had per-session volume since Vista, macOS needs a paid third-party app, and Linux has it if you know which PipeWire property to poke. None of them lets you set it from another room, from a script, or on a schedule.
Multi-room audio is either expensive or closed. The hardware that does it well costs more than the mini PC and the small board computer already sitting in the house, and it will not play a stream that came from an application on a laptop.
Assistants cannot drive audio properly. Give a model a volume API and it loops calls to fake a fade, and the result stutters, because a model round trip takes a second or two with jitter while an audio ramp needs a tick every 20 milliseconds.
The idea that makes it work: clients send intent, the hub owns the clock
The obvious design is a volume setter plus a client that ramps it, and it is unusable from anything except a local interface. Forty calls arriving a second apart produce audible stepping.
So every setter in the API takes a fade duration, and there are no relative setters at all. No volume up, no increments. Levels and targets are always absolute. One ticker inside the hub drives every in-flight ramp on an equal-power curve, and a new fade on a target supersedes the old one from wherever it currently sits, so nothing jumps.
The same reasoning makes handoff, broadcast, alarm, and scene application single atomic calls that return a job id, rather than sequences a client has to choreograph. An assistant cannot get the timing wrong if it never holds the timing.
Architecture
A hub and agent split communicating over gRPC. The hub owns the graph, coordinates zones, serves REST, WebSocket, and MCP, and supervises the sync server. An agent runs on each host, exposing that machine’s audio devices and per-app streams.
The platform layer carries three backends: PipeWire on Linux, Core Audio on macOS through a Swift helper, and WASAPI on Windows. A separate simulator crate implements the agent contract with virtual devices, so the entire hub is testable with no audio hardware present anywhere.
Sample-accurate playback across rooms is the hardest problem here and the most patent-dense, since synchronized group playback has been litigated heavily. Snapcast already solves it to about a millisecond, so SoundA2Z runs it as a supervised process and deliberately never links it. That boundary is a licensing decision written down as an architecture record, not a technical accident.
macOS: the platform that fought back hardest
Core Audio has no first-party concept of per-application enumeration, so telling the hub what a Mac is actually playing meant writing a small Swift helper that reads process identity a different way than a PID, since a PID recycles under exactly the conditions that make it worth watching in the first place.
A zone used to be a volume knob the hub set from a distance. It is now a supervised snapclient: the hub asks a sink-capable agent to run its own zones’ clients, and an orphaned one is provably killed by the same signal the test asserts, rather than assumed dead by absence of evidence.
It now ships as one signed .app that carries and supervises the hub itself, with a print-plan mode that states what a run would do without doing any of it, because a supervisor for other people’s processes should never be the thing you cannot trust to describe its own plan honestly.
Interested in the rest?
The full history, the architecture decisions, and the code behind this are available on request.