Films
Short films about Auracle: what it is, how to play it, and how it works underneath. Captions are on by default, and every film's full transcript is printed under it.
Everything you hear in them is Auracle: the music is scored for its own voices and played by its engine. The narration is synthetic (Kokoro-82M, offline).
Start here
Auracle 1:38
- 0:00 The problem
- 0:12 Auracle
- 0:19 Two patches, one pick
- 0:25 Real circuits
- 0:34 Playing it
- 0:46 Offers
- 0:58 Underneath
- 1:14 Every note
- 1:34 Play it
Transcript
Every synthesizer has a sound in it that's yours. Finding it means turning hundreds of knobs, one at a time. Auracle is a synthesizer that searches for your sound. It plays you two patches. You pick the one you like better. Every pick teaches it your taste, and it grows new patches toward it. Not samples. Real modular circuits, built and wired from scratch. Then you play it. Turn Bright, and it finds the knobs that make this patch brighter. Let it wander, and the knobs turn themselves toward your taste. Press Offer, and a new version grows from the sound in your hands. Blend into it. Take it, or pass. Either way, it learns. Open the circuit any time. Every knob is real, and you can watch the performance turn them. Underneath, a model of your taste bets on every choice before you make it, and keeps score in public. Every note in this film is Auracle. Free, open source, and running in your browser. Play it today.
How it works
How Auracle learns what you like 1:49
- 0:00 Choosing, not describing
- 0:11 What it listens for
- 0:28 A pick is evidence
- 0:38 Every taste that still fits
- 0:47 More than one taste
- 0:58 Forecasts, scored
- 1:11 The search
- 1:24 Reading what it learned
- 1:33 Learning while you play
- 1:40 In the open
Transcript
You know which of two sounds you like, long before you can say why. So Auracle never asks you to describe a sound. It asks you to choose. Behind every choice, it listens for the things you hear: how bright a sound is, how noisy, how it starts, and how fast it moves. That's eighteen measurements of every patch's sound, all from the same short phrase, and twenty-six more of how the patch is built. Each pick is evidence: you liked this set of measurements more than that one. Many tastes could explain one pick. A few picks rule most of them out. Auracle keeps every taste that still fits, weighted by how well it fits. With every answer, that cloud of possible tastes draws tighter. And taste isn't one direction. You can love dark drones and bright plucks. So the model keeps several lenses, and a sound only has to please one of them. Before every duel, it writes down a forecast. Afterwards, it checks. The TRUST view shows how honest those forecasts have been, even when the answer is: no better than a coin flip, yet. Then it searches. Evolution proposes new patches from a grammar of modules, and your taste tilts every proposal toward what you'll like. Thousands are heard silently. Only the best few ever reach you. You can read what it learned: a map of every patch you've heard, the styles it found, and each direction with its uncertainty. And when you play, it keeps listening. An offer you hear, then take or pass, counts just like a duel. Your taste, learned in the open. And it never leaves your browser.
Under the hood 2:17
- 0:00 Five crates
- 0:11 The genome
- 0:23 Compiling to DSP
- 0:31 The audition
- 0:40 Features
- 0:55 Utility
- 1:15 Calibration
- 1:24 Search
- 1:39 PERFORM's wiring
- 1:56 The runtime
- 2:06 Read it, run it
Transcript
Auracle is a Rust workspace of five crates, compiled to WebAssembly. Here's how a patch becomes a sound, and a choice becomes a model. A patch is a term in a typed grammar: a probabilistic program over modules. Every knob and every structural choice has a trace address, so a whole patch is one draw from a prior. It compiles to a quiver signal graph, which runs one sample at a time with no allocation on the audio path. Every candidate plays the same standard phrase, normalized for loudness, through a vetting gate that rejects silence, clipping and DC. From that phrase come eighteen perceptual features, from brightness and noisiness to envelope shape and three bands of modulation rate. Twenty-six structural ones come from the patch itself. Every one is standardized. Taste is a utility: the maximum over a few linear experts on those features. A duel, a keep or a cut, a star rating: each has its own likelihood. The posterior is sampled by Markov chain Monte Carlo, and each new answer reweights those samples until a refit is due. Every duel is forecast before it's answered. Each forecast is scored with a proper scoring rule, separately for each kind of evidence. Search targets a Boltzmann distribution: the grammar's prior, tilted by expected utility. Refinement is Metropolis-Hastings on the trace, through fugue-evo. A lock is exact conditioning. PERFORM's named controls are fixed directions in that standardized space of sound. For each patch, a finite-difference Jacobian and a ridge solve wire each control to its knobs. Every half of every control is then checked on real renders. In the browser, the engine runs in a worker, and the voices in an AudioWorklet. A render farm measures candidates in parallel. Every claim here has a measurement behind it, in the reference. Read it, run it, and change it.
The math 2:46
- 0:00 Intro
- 0:06 Utility
- 0:23 Likelihoods
- 0:44 Posterior
- 1:01 Calibration
- 1:17 Acquisition
- 1:31 Target
- 1:49 Refine
- 2:03 Locks
- 2:14 Perform
- 2:37 Outro
Transcript
The math inside Auracle, and why each piece has its shape. Each patch becomes forty-four features, standardized to one scale. Utility is the maximum over a few linear experts, never their average. So you can love dark drones and bright plucks, and each is scored by its own best lens. Three kinds of answer feed that one utility. A duel is Bradley-Terry, logistic in the utility difference. Cutting a patch is a kill, judged against a bar fitted per session. A picky day moves the bar. Not the taste. Stars fall between fitted cutpoints, so a harsh rater moves the cutpoints. With no hidden lens labels, every parameter is a real number. So plain Metropolis-Hastings fits it, and keeps five hundred draws. Between fits, each answer reweights the draws, exactly and nearly free. When the weights collapse, it pays for a refit. Each duel is forecast before you answer, then scored by Brier. Brier is a proper rule, so only an honest probability scores best. Accuracy cannot see overconfidence. Each kind of evidence gets its own score. Which pair should it ask about? Picking the most informative pair only tied random pairs. Thompson sampling lost. So pairs are random, and every duel is also an unbiased check. Search aims at a Boltzmann target, the grammar's prior times the exponential of beta times expected utility. The prior supplies parsimony as a probability, not a penalty to tune. Beta, at two, is the one dial between browsing and optimizing. Refinement is Metropolis-Hastings on the trace, through fugue-evo. It walks forty steps from each of the ten best patches. Keeping where each walk ends climbs the target instead of sampling it, which suits a shortlist. A lock is exact conditioning. Moves that change, delete or create a locked address are rejected. Checking births as well as deaths keeps detailed balance. PERFORM's controls are fixed directions in standardized sound. Bright is centroid plus rolloff. Each patch gets its own Jacobian from one nudged render per knob, since knobs act differently in each. A ridge solve picks at most four knobs for each control. Each half is then rendered for real, and closes if it stops moving the right way. Every constant here is in the reference, with its measurement where there is one.
The sound engine 2:49
- 0:00 Intro
- 0:07 Graph
- 0:23 Modules
- 0:41 Compile
- 0:54 Phrase
- 1:12 Vetting
- 1:23 Loudness
- 1:39 Features
- 2:10 Live
- 2:29 Farm
- 2:39 Outro
Transcript
This is how Auracle makes sound, from the patch graph to the live voices. Underneath is quiver, a modular synthesis library in Rust. On each tick, one sample moves through the whole graph. Continuous knobs are atomic values the audio thread reads, so turning one needs no recompile. The palette has forty-two modules, from a plucked string to sidechained dynamics. The filter is a state variable design, or a diode ladder that saturates harder one way. Audio and modulation are separate Rust types, so a mistyped patch cannot even be built. Every voice ends with a DC blocker where needed, an exponential envelope, and a limiter. Resonance and feedback are capped, so filters cannot oscillate and delays cannot run away. For comparison, every patch plays the same five second phrase. It holds a C, stabs an octave higher, and plays a two note chord. It ends on a low C, with a long release. The random seed is reset every render, so the samples repeat bit for bit. First, each raw render goes through a gate. It fails silence, runaway peaks, and signals dominated by DC. What fails is never played. Loudness is measured the broadcast way, with K weighting and gated blocks. Each patch is matched to minus eighteen loudness units, since louder wins comparisons. The gain stops short of clipping instead of limiting, so timbre is untouched. From that render come eighteen audio features. Four measure the spectrum's brightness and its movement, on a logarithmic frequency axis. Texture, level and envelope take seven more, and one measures the bass. Three are read from single notes. Three more are bands of motion on the held note. The bands run from half a hertz to two, and from two to eight. The fastest runs from eight to thirty. Twenty-six structural features come from the patch, with no render. Live, the same compiler builds four voices inside an AudioWorklet. Four more play PERFORM's offers, crossfaded at equal power. In steady state, the audio thread allocates nothing. A patch change fades out, rebuilds the voices in silence, and carries your held notes across. Auditions render in parallel, on up to six workers. Draws are indexed and absorbed in order, so the pool is identical at any width. One compiler serves search and stage, so what you play is what the model measured.