Density Engine for Perceptual Transparency & Hierarchy
We spent two years solving a problem that doesn't have a name yet.
Put it on one mix. You'll hear it in seconds.
You know the feeling: every channel is technically clean, the frequency balance checks out, the levels are right. But the mix still sounds congested and flat. You reach for another EQ, another compressor, another round of surgical fixes. None of it helps, because you're trying to fix channels when the problem is between them.
Dozens of sources all moving, modulating, and shifting at once. Your listener's brain is trying to track all of them, and it's exhausting. The congestion you're (unsuccessfully) trying to EQ your way out of? It's a bottleneck of perceptual density. Too many things demanding attention at once.
The mixes you admire have one thing in common: they have very few things fighting for your attention at any given moment. Listen to any mix you love. Count how few things are actually demanding your attention at any given time. The listener relaxes, and everything falls into place.
We don't really have a good elevator pitch for DEPTH. It doesn't have a frequency curve to show you or a gain reduction meter to point at. The closest we can get: it works on how your listener's brain decides what's in front and what's behind.
Your brain figures out what to focus on based on movement. Things that change rapidly grab attention. Things that are steady get filed as space. The room the music lives in. This is established psychoacoustics. It builds on six decades of neuroscience and studies on how your brain perceives sound. We didn't discover any of it, but we spent countless hours translating it into real, usable code.
DEPTH is session-wide. Every instance talks to every other instance in real time, so the plugin has a picture of your entire mix, not just one channel. No single-channel plugin knows what the rest of your session is doing. That's the fundamental limitation DEPTH was built to solve.
Auditory scene analysis, perceptual masking, attentional salience, spectral flux modeling. If you want to go deep, there are 60 cited papers at the bottom of this page.
Every instance shares its real-time spectral behavior with every other instance across your session.
Each element gets continuously ranked on a foreground-to-background gradient: how much spectral movement, how transient-dense, how much energy in the presence band.
Background elements receive subtle corrections. Their micro-level spectral movement is gently calmed so the ear reads them as environment instead of competing events. The sound is fully intact, but the brain places it differently.
Your listener goes from tracking twenty competing objects to focusing on three or four. Everything else becomes width, warmth, dimension. Same energy, same frequencies — but the mix opens up.
Background elements stop reading as separate objects and start filling out the stereo field. No widening plugin involved.
Multiple bass sources stop micro-competing and settle into one foundation. Most mud is a density problem, not a frequency problem.
The vocal doesn't get louder. Everything else just backs off the same perceptual space. Less competition, more presence.
Less congestion means your limiter works easier. Same LUFS, more punch.
We're aware this sounds like magic, so let's be specific about what the plugin does not do.
No EQ. Nothing is cut or boosted — your frequency balance stays where you put it. No compression. No threshold, no ratio, no dynamics processing. No stereo manipulation — no widening, no mid/side. The width you hear is perceptual. No phase or time tricks. Gain-domain only. And no AI. DEPTH doesn't have an opinion about your mix. It follows established psychoacoustic models, not machine learning.
Vocals, drums, bass, synths. One instance per bus, last in the chain.
DEPTH finds every other instance in your session automatically. No routing, no setup.
Set your separation and priority. Turn it up until your mix opens. That's the entire workflow.
The honest, plain-language answers for engineers who want to look under the hood.
Working as intended.
Each correction DEPTH makes is tiny on its own — well below the threshold where you'd notice it on a solo'd channel. But across 8 or 12 buses simultaneously, those micro-adjustments add up to a significant shift in how the full mix reads. More width, more front-to-back separation, less fatigue.
Try this: solo any track with DEPTH active. You'll hear almost nothing. Unsolo. The difference is immediate. That's the whole point. If you could hear it on solo, it would be coloring your sound.
Each instance continuously measures its own track: energy distribution, spectral movement, transient density, modulation behavior. That data gets shared across all instances, so the engine has a real-time picture of the full session.
From there it identifies where tracks are perceptually colliding and computes per-track correction curves: micro-level gain adjustments across the spectrum, thousands per second, each one clamped to stay below audibility on solo. The cumulative result across all buses: less crowding, clearer hierarchy, same tone and character.
The simplified version looks like this:
That curve is built from two weighted terms:
That formula is the simplified output. What feeds into it is where the complexity lives.
Each instance continuously measures 20+ psychoacoustic features: spectral flux, onset density, modulation depth, harmonic coherence, envelope regularity, presence-band energy, transient salience, among others. These aren't measured in isolation. They're cross-referenced against every other active instance, weighted according to established models for auditory scene analysis, masking, and attentional salience.
The math behind those two terms runs to 200+ lines of continuous operations per track, per frequency bin, per frame. The result is a correction curve that stays below JND thresholds on any individual channel but shifts the full mix noticeably when applied across the session. Everything is then clamped by your Max Intervention ceiling as a final safety net.
The research behind each of these measurements is cited at the bottom of this page.
Your ear is a contrast detector. It doesn't hear absolute values. It hears relationships. So when you shift the spectral behavior of 10 tracks by a tiny amount each, the relative change between them is significant, even though no individual track sounds different.
This is also why no single-channel plugin can get here. An EQ on your vocal has no idea what your pads are doing. DEPTH does.
Separation controls how much DEPTH differentiates between tracks. Higher values push background elements to behave more steadily, making them easier to distinguish from foreground events.
Priority controls how hard DEPTH pushes the foreground/background hierarchy. Higher values give your lead elements more room by pulling competing sources further back.
Max Intervention is the safety ceiling. The maximum amount any single track can be modified, in dB. Default is conservative. Most people never touch it.
No. Gain-domain only. No phase rotation, no time shifting, no pitch manipulation, no decorrelation.
All corrections are spectral gain adjustments: subtle changes to how energy is distributed across the frequency spectrum over time. Phase relationships and timing stay untouched.
We could get stronger separation with phase tricks, but the tradeoff is artifacts. We chose transparency.
Last insert on each bus, after everything else. DEPTH needs to see your final signal post-EQ, post-compression, post-saturation to make accurate decisions.
Phase alignment tools (like Sound Radix Pi) should go before DEPTH. It needs to see the signal as it will actually hit the mix bus.
We recommend buses rather than individual channels: vocal bus, drum bus, bass bus, synth bus, effects bus. Matches how your ear groups things and keeps the CPU reasonable.
The best explanation is the one you'll hear.
Try it on a mix.
14-day free trial included with both options.
VST3 · AU
DEPTH's processing engine is grounded in peer-reviewed psychoacoustic and cognitive research spanning six decades.
Every measurement, weighting, and perceptual threshold in DEPTH traces back to established research in auditory scene analysis, psychoacoustics, and cognitive neuroscience. The following bibliography represents the scientific foundation underlying DEPTH's 200+ lines of mathematical operations and 20+ measured psychoacoustic phenomena.
The foundational framework for how the brain separates complex sound mixtures into individual perceptual objects.
DEPTH measures spectral flux per instance as a primary input to its behavioral classification engine.
Simultaneous and temporal masking models inform where cross-track intervention is needed.
DEPTH's priority engine models which elements will capture the listener's attention.
Modulation rate and coherence are primary grouping cues. Elements with similar modulation patterns fuse perceptually.
Rapid onsets capture attention; steady-state signals are parsed as background.
Harmonically related partials fuse into a single percept. DEPTH measures coherence to determine source membership.
Steady envelopes read as background; irregular envelopes read as events demanding attention.
The frequency region of peak sensitivity and speech intelligibility, weighted specifically in DEPTH's priority calculations.
All corrections are clamped to stay below perceptual JND limits per frequency band and per track.
DEPTH's gain-domain processing is informed by how the auditory system itself uses gain control to adapt to stimulus statistics.
The core problem DEPTH addresses: multiple competing objects degrade perception even without spectral overlap.
The perceptual effect DEPTH produces: foreground/background hierarchy and perceived spatial depth.