OBSIDIAN

DEPTH

Density Engine for Perceptual Transparency & Hierarchy

We spent two years solving a problem that doesn't have a name yet.
Put it on one mix. You'll hear it in seconds.

LO HI
DENSITY · BEFORE
The problem nobody talks about

There's a reason your mix sounds flat, and it has nothing to do with EQ.

You know the feeling: every channel is technically clean, the frequency balance checks out, the levels are right. But the mix still sounds congested and flat. You reach for another EQ, another compressor, another round of surgical fixes. None of it helps, because you're trying to fix channels when the problem is between them.

Dozens of sources all moving, modulating, and shifting at once. Your listener's brain is trying to track all of them, and it's exhausting. The congestion you're (unsuccessfully) trying to EQ your way out of? It's a bottleneck of perceptual density. Too many things demanding attention at once.

The mixes you admire have one thing in common: they have very few things fighting for your attention at any given moment. Listen to any mix you love. Count how few things are actually demanding your attention at any given time. The listener relaxes, and everything falls into place.

Not an EQ · Not a compressor · A perceptual engine

Something between a mix tool and a psychoacoustic model. Here's our best attempt at explaining it.

We don't really have a good elevator pitch for DEPTH. It doesn't have a frequency curve to show you or a gain reduction meter to point at. The closest we can get: it works on how your listener's brain decides what's in front and what's behind.

Your brain figures out what to focus on based on movement. Things that change rapidly grab attention. Things that are steady get filed as space. The room the music lives in. This is established psychoacoustics. It builds on six decades of neuroscience and studies on how your brain perceives sound. We didn't discover any of it, but we spent countless hours translating it into real, usable code.

DEPTH is session-wide. Every instance talks to every other instance in real time, so the plugin has a picture of your entire mix, not just one channel. No single-channel plugin knows what the rest of your session is doing. That's the fundamental limitation DEPTH was built to solve.

Auditory scene analysis, perceptual masking, attentional salience, spectral flux modeling. If you want to go deep, there are 60 cited papers at the bottom of this page.

01

Cross-instance analysis

Every instance shares its real-time spectral behavior with every other instance across your session.

02

Behavioral classification

Each element gets continuously ranked on a foreground-to-background gradient: how much spectral movement, how transient-dense, how much energy in the presence band.

03

Perceptual shaping

Background elements receive subtle corrections. Their micro-level spectral movement is gently calmed so the ear reads them as environment instead of competing events. The sound is fully intact, but the brain places it differently.

04

Depth emerges

Your listener goes from tracking twenty competing objects to focusing on three or four. Everything else becomes width, warmth, dimension. Same energy, same frequencies — but the mix opens up.

What you'll hear

The difference between a good mix and a mix that sounds like it cost $10,000.

Wider soundstage

Background elements stop reading as separate objects and start filling out the stereo field. No widening plugin involved.

Warmer low end

Multiple bass sources stop micro-competing and settle into one foundation. Most mud is a density problem, not a frequency problem.

Vocal clarity

The vocal doesn't get louder. Everything else just backs off the same perceptual space. Less competition, more presence.

More headroom

Less congestion means your limiter works easier. Same LUFS, more punch.

Skeptical? Good

What DEPTH isn't

We're aware this sounds like magic, so let's be specific about what the plugin does not do.

No EQ. Nothing is cut or boosted — your frequency balance stays where you put it. No compression. No threshold, no ratio, no dynamics processing. No stereo manipulation — no widening, no mid/side. The width you hear is perceptual. No phase or time tricks. Gain-domain only. And no AI. DEPTH doesn't have an opinion about your mix. It follows established psychoacoustic models, not machine learning.

Setup

Three steps. No presets. Your mix is the preset.

1

Insert on each bus

Vocals, drums, bass, synths. One instance per bus, last in the chain.

2

Instances auto-connect

DEPTH finds every other instance in your session automatically. No routing, no setup.

3

Adjust to taste

Set your separation and priority. Turn it up until your mix opens. That's the entire workflow.

FAQ

How does DEPTH actually work?

The honest, plain-language answers for engineers who want to look under the hood.

The big question I hear a big mix difference, but I can't hear an effect on solo. What's going on?

Working as intended.

Each correction DEPTH makes is tiny on its own — well below the threshold where you'd notice it on a solo'd channel. But across 8 or 12 buses simultaneously, those micro-adjustments add up to a significant shift in how the full mix reads. More width, more front-to-back separation, less fatigue.

Try this: solo any track with DEPTH active. You'll hear almost nothing. Unsolo. The difference is immediate. That's the whole point. If you could hear it on solo, it would be coloring your sound.

Under the hood What is DEPTH actually doing to my audio?

Each instance continuously measures its own track: energy distribution, spectral movement, transient density, modulation behavior. That data gets shared across all instances, so the engine has a real-time picture of the full session.

From there it identifies where tracks are perceptually colliding and computes per-track correction curves: micro-level gain adjustments across the spectrum, thousands per second, each one clamped to stay below audibility on solo. The cumulative result across all buses: less crowding, clearer hierarchy, same tone and character.

The math Can you explain the processing more precisely?

The simplified version looks like this:

output = input × (1 + tiny_time_varying_curve)

That curve is built from two weighted terms:

tiny_curve = separation_amount × separation_targets
            + priority_amount × priority_targets

where separation_targets and priority_targets are each derived from
200+ lines of mathematical operations measuring and cross-referencing
20+ individual psychoacoustic phenomena in real time

That formula is the simplified output. What feeds into it is where the complexity lives.

Each instance continuously measures 20+ psychoacoustic features: spectral flux, onset density, modulation depth, harmonic coherence, envelope regularity, presence-band energy, transient salience, among others. These aren't measured in isolation. They're cross-referenced against every other active instance, weighted according to established models for auditory scene analysis, masking, and attentional salience.

The math behind those two terms runs to 200+ lines of continuous operations per track, per frequency bin, per frame. The result is a correction curve that stays below JND thresholds on any individual channel but shifts the full mix noticeably when applied across the session. Everything is then clamped by your Max Intervention ceiling as a final safety net.

The research behind each of these measurements is cited at the bottom of this page.

Why it works Why do tiny changes add up to such a big difference?

Your ear is a contrast detector. It doesn't hear absolute values. It hears relationships. So when you shift the spectral behavior of 10 tracks by a tiny amount each, the relative change between them is significant, even though no individual track sounds different.

This is also why no single-channel plugin can get here. An EQ on your vocal has no idea what your pads are doing. DEPTH does.

The controls What do Separation and Priority actually do?

Separation controls how much DEPTH differentiates between tracks. Higher values push background elements to behave more steadily, making them easier to distinguish from foreground events.

Priority controls how hard DEPTH pushes the foreground/background hierarchy. Higher values give your lead elements more room by pulling competing sources further back.

Max Intervention is the safety ceiling. The maximum amount any single track can be modified, in dB. Default is conservative. Most people never touch it.

Technical detail Is DEPTH manipulating phase or time?

No. Gain-domain only. No phase rotation, no time shifting, no pitch manipulation, no decorrelation.

All corrections are spectral gain adjustments: subtle changes to how energy is distributed across the frequency spectrum over time. Phase relationships and timing stay untouched.

We could get stronger separation with phase tricks, but the tradeoff is artifacts. We chose transparency.

Setup Where should DEPTH go in my signal chain?

Last insert on each bus, after everything else. DEPTH needs to see your final signal post-EQ, post-compression, post-saturation to make accurate decisions.

Phase alignment tools (like Sound Radix Pi) should go before DEPTH. It needs to see the signal as it will actually hit the mix bus.

We recommend buses rather than individual channels: vocal bus, drum bus, bass bus, synth bus, effects bus. Matches how your ear groups things and keeps the CPU reasonable.

The best explanation is the one you'll hear.
Try it on a mix.

Launch pricing
€179 €229

One-time purchase. Free updates for life.

Coming Soon
Rent to own
€19 /mo

12 months. Then it's yours forever.

Coming Soon

14-day free trial included with both options.

VST3 · AU

60 Cited Papers · View Research Foundations

DEPTH's processing engine is grounded in peer-reviewed psychoacoustic and cognitive research spanning six decades.

Every measurement, weighting, and perceptual threshold in DEPTH traces back to established research in auditory scene analysis, psychoacoustics, and cognitive neuroscience. The following bibliography represents the scientific foundation underlying DEPTH's 200+ lines of mathematical operations and 20+ measured psychoacoustic phenomena.

Auditory Scene Analysis

The foundational framework for how the brain separates complex sound mixtures into individual perceptual objects.

  1. Bregman, A. S. (1990). Auditory Scene Analysis: The Perceptual Organization of Sound. MIT Press.
  2. Bregman, A. S. & McAdams, S. (1994). Auditory Scene Analysis. In International Encyclopedia of the Social and Behavioral Sciences. Pergamon.
  3. Carlyon, R. P. (2004). How the brain separates sounds. Trends in Cognitive Sciences, 8(10), 465–471.
  4. Ciocca, V. (2008). The auditory organization of complex sounds. Frontiers in Bioscience, 13, 148–169.
  5. Sussman, E. S. (2017). Auditory Scene Analysis: An Attention Perspective. Journal of Speech, Language, and Hearing Research, 60(10), 2913–2926.

Spectral Flux Perception

DEPTH measures spectral flux per instance as a primary input to its behavioral classification engine.

  1. Weineck, K. et al. (2022). Neural synchronization is strongest to the spectral flux of slow music and depends on familiarity and beat salience. eLife, 11, e75515.
  2. Teki, S. et al. (2011). Brain bases for auditory stimulus-driven figure-ground segregation. Journal of Neuroscience, 31(1), 164–171.
  3. Elhilali, M. et al. (2009). A spectro-temporal modulation index for assessment of speech intelligibility. Speech Communication, 51(12), 1108–1124.

Perceptual Masking

Simultaneous and temporal masking models inform where cross-track intervention is needed.

  1. Moore, B. C. J. (2012). An Introduction to the Psychology of Hearing (6th ed.). Brill.
  2. Moore, B. C. J. (2007). Psychoacoustics. In Rossing, T. (Ed.), Springer Handbook of Acoustics. Springer.
  3. Moore, B. C. J. (1985). Additivity of simultaneous masking, revisited. Journal of the Acoustical Society of America, 78(2), 488–494.
  4. Moore, B. C. J. & Oxenham, A. J. (1998). Psychoacoustic consequences of compression in the peripheral auditory system. Psychological Review, 105(1), 108–124.
  5. Moore, B. C. J. (2010). Masking in the human auditory system. In The Oxford Handbook of Auditory Science, Vol. 3. Oxford University Press.
  6. Patterson, R. D. (1976). Auditory filter shapes derived with noise stimuli. Journal of the Acoustical Society of America, 59(3), 640–654.

Attentional Salience & Selective Attention

DEPTH's priority engine models which elements will capture the listener's attention.

  1. Shinn-Cunningham, B. G. (2008). Object-based auditory and visual attention. Trends in Cognitive Sciences, 12(5), 182–186.
  2. Lakatos, P. et al. (2013). The spectrotemporal filter mechanism of auditory selective attention. Neuron, 77(4), 750–761.
  3. Gutschalk, A. & Dykstra, A. R. (2014). Functional imaging of auditory scene analysis. Hearing Research, 307, 98–110.
  4. Rimmele, J. M. et al. (2015). The role of prediction in auditory scene analysis. Frontiers in Neuroscience, 9, 143.
  5. Mesgarani, N. & Chang, E. F. (2012). Selective cortical representation of attended speaker in multi-talker speech perception. Nature, 485, 233–236.

Modulation Depth & AM/FM Coherence

Modulation rate and coherence are primary grouping cues. Elements with similar modulation patterns fuse perceptually.

  1. McAdams, S. (1989). Segregation of concurrent sounds. I: Effects of frequency modulation coherence. Journal of the Acoustical Society of America, 86(6), 2148–2159.
  2. Marin, C. M. H. & McAdams, S. (1991). Segregation of concurrent sounds. II: Effects of spectral envelope tracing, FM coherence, and FM width. J. Acoust. Soc. Am., 89(1), 341–351.
  3. Culling, J. F. & Summerfield, Q. (1995). The role of frequency modulation in the perceptual segregation of concurrent vowels. J. Acoust. Soc. Am., 98(2), 837–846.
  4. Bregman, A. S., Levitan, R. & Liao, C. (1990). Fusion of auditory components: Effects of the frequency of amplitude modulation. Perception & Psychophysics, 47, 68–73.
  5. Summerfield, Q. & Culling, J. F. (1992). Auditory segregation of competing voices: absence of effects of FM or AM coherence. Phil. Trans. R. Soc. Lond. B, 336, 357–366.

Onset Density & Transient Salience

Rapid onsets capture attention; steady-state signals are parsed as background.

  1. Bregman, A. S. & Doehring, P. (1984). Fusion of simultaneous tonal glides: The role of parallelness and simple frequency relations. Perception & Psychophysics, 36, 251–256.
  2. Darwin, C. J. & Carlyon, R. P. (1995). Auditory grouping. In Moore, B. C. J. (Ed.), Handbook of Perception and Cognition: Hearing (2nd ed., pp. 387–424). Academic Press.
  3. Cusack, R. & Carlyon, R. P. (2003). Perceptual asymmetries in audition. J. Exp. Psychol.: Hum. Percept. Perform., 29(3), 713–725.
  4. Nelken, I. (2004). Processing of complex stimuli and natural scenes in the auditory cortex. Current Opinion in Neurobiology, 14(4), 474–480.

Harmonic Coherence & Pitch-Based Grouping

Harmonically related partials fuse into a single percept. DEPTH measures coherence to determine source membership.

  1. Moore, B. C. J., Glasberg, B. R. & Peters, R. W. (1986). Thresholds for hearing mistuned partials as separate tones in harmonic complexes. J. Acoust. Soc. Am., 80, 479–483.
  2. Chalikia, M. H. & Bregman, A. S. (1993). The perceptual segregation of simultaneous vowels with harmonic, shifted, and random components. Percept. & Psychophys., 53, 125–133.
  3. de Cheveigné, A. (1993). Separation of concurrent harmonic sounds: Fundamental frequency estimation and a time-domain cancellation model. J. Acoust. Soc. Am., 93(6), 3271–3290.

Envelope Regularity & Temporal Fine Structure

Steady envelopes read as background; irregular envelopes read as events demanding attention.

  1. Joris, P. X., Schreiner, C. E. & Rees, A. (2004). Neural processing of amplitude-modulated sounds. Physiological Reviews, 84(2), 541–577.
  2. Moore, B. C. J. (2008). The role of temporal fine structure processing in pitch perception, masking, and speech perception. J. Assoc. Res. Otolaryngol., 9(4), 399–406.
  3. Ding, N. & Simon, J. Z. (2013). Adaptive temporal encoding leads to a background-insensitive cortical representation of speech. J. Neurosci., 33(13), 5728–5735.

Presence-Band Energy (2–5 kHz)

The frequency region of peak sensitivity and speech intelligibility, weighted specifically in DEPTH's priority calculations.

  1. Fletcher, H. & Munson, W. A. (1933). Loudness, its definition, measurement and calculation. J. Acoust. Soc. Am., 5(2), 82–108.
  2. Robinson, D. W. & Dadson, R. S. (1956). A re-determination of the equal-loudness relations for pure tones. British Journal of Applied Physics, 7, 166–181.
  3. ISO 226:2003. Acoustics — Equal-loudness-level contours.
  4. ANSI S3.5-1997. Methods for Calculation of the Speech Intelligibility Index.

Just-Noticeable Difference (JND) Thresholds

All corrections are clamped to stay below perceptual JND limits per frequency band and per track.

  1. Weber, E. H. (1834). De Pulsu, Resorptione, Auditu et Tactu: Annotationes Anatomicae et Physiologicae.
  2. Fastl, H. & Zwicker, E. (2007). Just-Noticeable Sound Changes. In Psychoacoustics: Facts and Models (3rd ed.). Springer.
  3. Hellman, R. P. et al. (1993). Just noticeable differences for intensity and their relation to loudness. J. Acoust. Soc. Am., 93(2), 908–918.
  4. Florentine, M. et al. (1993). Intensity JNDs at equal-loudness levels in normal and pathological ears. J. Acoust. Soc. Am., 93(5), 2790–2796.
  5. McShefferty, D., Whitmer, W. M. & Akeroyd, M. A. (2015). The just-noticeable difference in speech-to-noise ratio. Trends in Hearing, 19, 1–9.

Auditory Gain Control & Neural Adaptation

DEPTH's gain-domain processing is informed by how the auditory system itself uses gain control to adapt to stimulus statistics.

  1. Dean, I., Harper, N. S. & McAlpine, D. (2005). Neural population coding of sound level adapts to stimulus statistics. Nature Neuroscience, 8, 1684–1689.
  2. Rabinowitz, N. C. et al. (2011). Contrast gain control in auditory cortex. Neuron, 70(6), 1178–1191.
  3. Robinson, B. L. et al. (2016). Meta-adaptation in the auditory midbrain under cortical influence. Nature Communications, 7, 13442.
  4. Willmore, B. D. B. et al. (2014). Neural adaptation to stimulus statistics in the auditory system. J. Neurosci., 34(27), 9109–9119.
  5. Mesgarani, N. et al. (2014). Phonetic feature encoding in human superior temporal gyrus. Science, 343(6174), 1006–1010.

Informational Masking & Perceptual Crowding

The core problem DEPTH addresses: multiple competing objects degrade perception even without spectral overlap.

  1. Ihlefeld, A. & Shinn-Cunningham, B. (2008). Spatial release from energetic and informational masking in a divided attention task. J. Acoust. Soc. Am., 123(6), 4380–4392.
  2. Kidd, G. et al. (2008). Informational masking. In Yost, W. A. et al. (Eds.), Auditory Perception of Sound Sources (pp. 143–189). Springer.
  3. Durlach, N. I. et al. (2003). Note on informational masking. J. Acoust. Soc. Am., 113(6), 2984–2987.
  4. Shinn-Cunningham, B. G. & Best, V. (2008). Selective attention in normal and impaired hearing. Trends in Amplification, 12(4), 283–299.

Auditory Depth & Distance Perception

The perceptual effect DEPTH produces: foreground/background hierarchy and perceived spatial depth.

  1. Zahorik, P., Brungart, D. S. & Bronkhorst, A. W. (2005). Auditory distance perception in humans: A summary of past and present research. Acta Acustica united with Acustica, 91(3), 409–420.
  2. Kopčo, N. & Shinn-Cunningham, B. G. (2011). Effect of stimulus spectrum on distance perception for nearby sources. J. Acoust. Soc. Am., 130(3), 1530–1541.

Comprehensive Textbooks

  1. Bregman, A. S. (1990). Auditory Scene Analysis. MIT Press.
  2. Moore, B. C. J. (2012). An Introduction to the Psychology of Hearing (6th ed.). Brill.
  3. Fastl, H. & Zwicker, E. (2007). Psychoacoustics: Facts and Models (3rd ed.). Springer.
  4. Gelfand, S. A. (2017). Hearing: An Introduction to Psychological and Physiological Acoustics (6th ed.). CRC Press.
  5. Plack, C. J. (2014). The Sense of Hearing (2nd ed.). Psychology Press.
  6. Yost, W. A. (2006). Fundamentals of Hearing: An Introduction (5th ed.). Academic Press.