Claiming that a 10db pin drop can be accurately reproduced with only 3 bits per sample without dithering is a rather extraordinary claim, given that human hearing is sensitive to far more than 8 levels of amplitude <= 10db absolute SPL.
A pin drop contains a continuous and gradual decay of the ringing over time, which most certainly is audible (at least subjectively) in more gradations than just 8 levels before reaching 0.
But this isn’t just subjective: Studies have confirmed that humans can hear decibel differences of 1db quite reliably, and can even hear as little as 0.2db subconsciously (this can be objectively measured)!
So long as you can hear a pin drop ringing at 1db, and also at 10db, it’s therefore obviously true that 3 bits per sample (undithered) is insufficient to express this (keep in mind, these samples are linear amplitudes!)
So I’m not sure how you can possibly claim that 3bits is enough to replicate the continuous amplitude decay of a pin drop’s ringing sound. What am I missing?
That’s said, I will now go read your link. I agree dithering is one viable way to expand the dynamic range, but you also seem to be claiming this 3bit pin drop is true without dithering.
I may be misunderstanding some of the math, but it's important to keep in mind that decibels are logarithmic. Adding 6 dB doubles the amplitude, which is why adding another bit to a digital audio format adds 6 dB of dynamic range.
Since we can't hear any sound quieter than -8 dB SPL, then (I think) we can't hear a difference smaller than the 6 dB between -8 and -2 dB because that difference is quieter than the threshold of human perception. Similarly, the next perceptible change in amplitude would be to +1.5 dB (3x the threshold of perception), then 4 dB (4x), 6 dB (5x), 7.6 dB (6x), 8.9 dB (7x), and 10 dB (8x).
So, since 10 dB is 8 times the threshold of human perception, I think eight discrete values should be enough to represent everything we can hear up to that point.
Of course, I'm assuming we can only hear in discrete increments of the minimum threshold of perception, -- it makes sense to me that we would be unable to hear a difference in amplitude quieter than the threshold -- but this could very well be a mistaken assumption. (It's also worth pointing out that 10 dB is very, very quiet -- the article I linked says 20 dB is the standard noise floor for a soundproof room/recording studio.)
Another way of thinking about this is in terms of quantization error. Our 3-bit pin drop signal can be interpreted as the original analog signal plus a quantization error, and that error is clearly always quieter than the minimum threshold of perception.
Edit: This article [0] contains a table for the minimum detectable difference in sound level. It shows that, at a level of 5
dB SPL, the just-noticeable difference in amplitude is 2.5 dB at the frequency range of peak sensitivity -- which is pretty close to the numbers I found based on discrete multiples of the absolute threshold. So it looks like 3 bits is indeed sufficient for our pin drop (unless I'm mistaken somewhere).
Firstly: I agree that with dithering, properly mastered CD quality should be more than enough in almost all cases. Yet, despite my agreement here, the meta-analysis I posted above clearly demonstrates that empirically we are both wrong! Obviously, it's bad to try to deny robust evidence just because our theories don't fit it; that's the opposite of scientific progress. Clearly, further investigation and possibly even theoretical adjustments will be required here, no matter how much we may believe we have had the math and psycho-acoustics nailed down.
I think there's a lot of subtlety being missed between dynamic range and resolution. For example, I think your JND (just noticeable difference) assumption is incorrect; my understanding of the scientific consensus is that a 1db difference at any volume is consciously perceivable to "normal" human hearing (and 3db to virtually all humans), and a 0.2db difference has been shown to be subconsciously perceivable by most. This places the required bit depth (not counting dithering) of an audio recording spanning all human perception WELL beyond 96db, and probably somewhere in excess of 120-140db! This is consensus, BTW; hence most CD-apologists people defer to the dithering argument :)
Also, I did finish reading that article up to the point about bit depth, and unfortunately there's quite a bit of either bad and misleading data in there unfortunately. While the article does contain some very good true info, it's sadly riddled with enough bad/outdated claims that I can't really recommend it to anyone as a reputable/trustworthy source of truth. It strikes me very much as if the author is unaware of how imprecise our approximate understanding of psycho-accoustics actually is, when it makes extremely overconfident claims like "CD quality will be enough FOREVER".
For example, the point made about near infrared being invisible to all humans is obviously false: most humans can't, but a certain percentage of blue-eyed humans can see near infra-red (including remote control IR). I've known one such person personally, and have tested and confirmed this thoroughly. I agree that it's silly to try to make a TV that emits these frequencies (and cameras that capture them), but my point here is simply: there's a lot of bad info in that article, perhaps shamefully so for an article making such bold and confident assertions under the name of "science".
Regarding bit depth, the article doesn't really even try to dispute the fact that 96db is insufficient; it just says with dithering, we shouldn't have to worry about it. I'd love to agree, but the combination of the meta-analysis that shows it's not sufficient is all I need to prove that your linked article is simply wrong.
There many be a wide range of reasons why the meta-analysis found high res audio to sound slightly better, but none of that invalidates the validity of the results themselves.
For example, maybe most DACs aren't very good at replicating dithered subtleties (I'm just speculating examples here) -- supposing such a common problem exists, then we could debate whether it makes more sense to improve DACs with fancy improvements to dither reconstruction filtering -- OR -- we could just encode at least 140db of dynamic range in the sample bit depth in the first place! The latter seems a far simpler and less-overengineered solution to me.
Dithering is fine when it works, but let's be honest: fundamentally, dithering is a kind of compression algorithm that encodes a greater dynamic range within a signal with an artificially constrained dynamic range (but with the side-effect that not all DACs will handle it equally well). In the modern digital world, there's no need to rely on dithering any more in audio signals, in the same way that we don't see dithered 256-color images much any more. Why don't we leave compression up to the compression algorithms, rather than promoting ACDs/DACs that are glorified compression/decompression hardware? Modern lossless compression algorithms are vastly better anyway, if bitrate savings is the goal.
> I think there's a lot of subtlety being missed between dynamic range and resolution. For example, I think your JND (just noticeable difference) assumption is incorrect; my understanding of the scientific consensus is that a 1db difference at any volume is consciously perceivable to "normal" human hearing (and 3db to virtually all humans), and a 0.2db difference has been shown to be subconsciously perceivable by most.
The Wikipedia article about JND seems to imply that the general 1 dB rule may not necessarily apply near the limits of perception. I haven't found a whole lot of more detailed information; however, the table I linked in the edit to my previous comment seems to indicate that the JND in dB gets much higher at very quiet volumes.
> For example, the point made about near infrared being invisible to all humans is obviously false: most humans can't, but a certain percentage of blue-eyed humans can see near infra-red (including remote control IR)
The article does mention that at some IR remotes can be visible:
> Can you see the Apple Remote's LED flash when you press a button? No? Not even the tiniest amount? Try a few other IR remotes; many use an IR wavelength a bit closer to the visible band, around 310-350THz. You won't be able to see them either. The rest emit right at the edge of visibility from 350-380 THz and may be just barely visible in complete blackness with dark-adjusted eyes. All would be blindingly, painfully bright if they were well inside the visible spectrum.
It has a footnote with more detailed information:
> The original version of this article stated that IR LEDs operate from 300-325THz (about 920-980nm), wavelengths that are invisible. Quite a few readers wrote to say that they could in fact just barely see the LEDs in some (or all) of their remotes. Several were kind enough to let me know which remotes these were, and I was able to test several on a spectrometer. Lo and behold, these remotes were using higher-frequency LEDs operating from 350-380THz (800-850nm), just overlapping the extreme edge of the visible range.
I don't have enough time or knowledge to really dive in to the scientific studies. That being said it's worth noting that a large number of studies that have found no perceptible difference between CD-quality and higher-resolution audio, and (as you mention) there are a lot of ways for minute hardware differences to skew the results either way.
> Dithering is a kind of compression algorithm that encodes a greater dynamic range within a signal with an artificially constrained dynamic range
Dithering does more than just expanding the dynamic range: it ensures that quantization error produces a predicable noise floor rather than input-dependent distortion. This video, by the author of the xiph.org article I linked before, has a good explanation and demonstration of dithering (use the chapter selection menu to skip to the "dither" section): https://xiph.org/video/vid2.shtml
A pin drop contains a continuous and gradual decay of the ringing over time, which most certainly is audible (at least subjectively) in more gradations than just 8 levels before reaching 0.
But this isn’t just subjective: Studies have confirmed that humans can hear decibel differences of 1db quite reliably, and can even hear as little as 0.2db subconsciously (this can be objectively measured)!
So long as you can hear a pin drop ringing at 1db, and also at 10db, it’s therefore obviously true that 3 bits per sample (undithered) is insufficient to express this (keep in mind, these samples are linear amplitudes!)
So I’m not sure how you can possibly claim that 3bits is enough to replicate the continuous amplitude decay of a pin drop’s ringing sound. What am I missing?
That’s said, I will now go read your link. I agree dithering is one viable way to expand the dynamic range, but you also seem to be claiming this 3bit pin drop is true without dithering.