The Quite OK Audio Format for Fast, Lossy Compression (opens in new tab)

(qoaformat.org)

140 pointssmlckz2y ago63 comments

63 comments

has anyone benchmarked qoa to see roughly how many instructions per sample it needs? all i see here is that it's more than adpcm and less than mp3, but those differ by orders of magnitude

like, can you reasonably qoa-compress real-time 16ksps audio on a 16 megahertz atmega328?

hmm, https://phoboslab.org/log/2023/04/qoa-specification has some benchmark results, let's see... seems like he encoded 9807 seconds of 44.1ksps stereo in 25.8 seconds and decoded it in 3.00 seconds on an i7-6700k running singlethreaded. what does that imply for other machines?

it seems to be integer code (because reproducibility between the predictor in encoding and decoding is important, and a significant part of it is 16-bit. https://ark.intel.com/content/www/xl/es/ark/products/88195/i... says it's a 4.2 gigahertz skylake. agner says skylake can do 4–6 ipc (well, μops/cycle) https://www.agner.org/optimize/blog/read.php?i=628, coincidentally testing on an i7-6700k himself, but let's assume it's 3 ipc, because it's usually hard to reach even that level of ilp in useful code

so that's about 380 μops per sample if i'm doing my math right; that might be on the order of 400 32-bit integer instructions per sample on an in-order processor. if (handwaving wildly now!) that's 600 8-bit instructions, the atmega328 should be able to encode somewhere in the range of 16–32 kilosamples per second

so, quite plausibly

for decoding the same math gives 43 μops per sample rather than 380

i'm very interested to hear anyone else's benchmarks or calculations

g0xA52A2A2y ago

Some previous discussions.

3 months ago - https://news.ycombinator.com/item?id=35738817

6 months ago - https://news.ycombinator.com/item?id=34625573

kragen2y ago

thank you very much

these had crucial information for me

mips_r4300i2y ago

Comparing against 4bit ADPCM, which is already able to give quite good performance as long as your sample rates are relatively modern, this only improves it to 3.2 bits. It is fast, but ADPCM is also fast.

Would be nice to see joint stereo support. If you were to take ADPCM or this OK format and try to encode any stereo music with it, you will need 2 channels. However, there is an extremely advantageous optimization that can be made here - most music is largely center panned, so both channels are almost the same. With joint stereo you record one channel (either by picking one or mixing to an average) and then you can store the difference for the other channel which will occupy a lot fewer bits, assuming you are able to quantize away the increased entropy.

For example, instead of using two 4bit ADPCM channels for stereo, which would only be a 50% savings over uncompressed, you could probably use an average of 5 bits per sample.

anotherhue2y ago

> Would be nice to see joint stereo support

This was/is available in MP3 since forever, so seems a reasonable request.

https://wiki.hydrogenaud.io/index.php?title=Intensity_stereo

gaazoh2y ago

I like the philosophy of QOA (and other similar projects, including QOI and TinyVG), but unlike others, it seems like it's not ready to use yet, see https://github.com/phoboslab/qoa/issues/25

> I have just pushed a workaround to master. [...]

> This still introduces audible artifacts when the weights reset. It prevents the LMS from exploding, but is far from perfect :/

This, combined with the fact that that issue is still open mean that a breaking change is still to be expected.

codeflo2y ago

It's interesting that this works in the time domain (instead of frequency domain), and I wonder what the resulting quality limitations are, if any. The sound samples on the demo page, at the least the dozen I clicked on, didn't seem all that challenging. Few, mostly synthesized instruments, low dynamic range. My ears aren't good enough to evaluate audio codecs anyway, however.

Pet_Ant2y ago

What is the LFE channel?

It should be spelled out explicitly, but I figured out the rest

L-Left,R-Right,C-Center,FL-Front Left,FR-FrontRight,SL-SideLeft,SR-SideRight,BL-BackLeft,BR-BackRight

---

Edit: LFE-LowFrequencyEffects... so subwoofer?

https://www.dolby.com/uploadedFiles/Assets/US/Doc/Profession...

ogurechny2y ago

LFE audio channel is different from subwoofer output.

Subwoofers come with multichannel audio systems in which directional speakers usually can't cover the lower range of audio frequencies. They are responsible for bass content from all channels, and get it from software or hardware crossover filter which is independent from specific input formats. Placement of low frequency speaker does not matter much because of human perception.

LFE track is an additional effects channel for movie theaters and similar amusement rides in which audio system plays low frequencies from other channels just fine. Dedicated LFE emitter then adds rattling and other wub-wub effects without overloading audio speakers with all that extra energy. Movies that lack car chases and explosions routinely have completely silent LFE tracks.

Doxin2y ago

So it's essentially a bass shaker track?

1 more reply

entropicdrifter2y ago

LFE is an industry standard term for the subwoofer channel. It's the ".1" in "5.1","6.1","7.1" etc

ok_dad2y ago

LFE is usually a bass shaker which is a subwoofer but it moves a weight instead of a cone, so you get vibrations in your seat. It stimulates movement to your body somewhat, I use two for my sim racing rig, one under my seat to inform me of the car dynamics and immersive feeling, one under my pedals to inform me when ABS is active and when my tires are spinning.

samplatt2y ago

LFE can mean "bass shaker", but it's an industry-standard term invented by Dolby that effectively means "between 3 and 120hz", which usually means "subwoofer".

These days crossover points are very configurable. Most bass shakers are rated for use between 20hz and 200hz.

bravura2y ago

Low frequency energy, I assume. Ie bass. Your subwoofer or „bottoms“ if you have several.

MobiusHorizons2y ago

Seems to have similar design criteria as opus but I don’t see any comparison.

Turing_Machine2y ago

I looked around, but didn't see any mention of potential patent issues. I assume that this has been considered? The Ogg Vorbis people spent a lot of time on that back when they were developing their format.

Other than that, looks great!

speedgoose2y ago

The website says it’s made in Hesse. No software patents to care about there.

https://en.m.wikipedia.org/wiki/Software_patents_under_the_E...

morelisp2y ago

Probably the most infamous audio format patent ever was owned by a German research institute.

1 more reply

IshKebab2y ago

You can absolutely patent software in Europe. Sorry. It's a common misconception that you can't. There's a stupid dance you have to do so it isn't technically "software" that you're patenting... but really it is.

1 more reply

Turing_Machine2y ago

Maybe not, but that doesn't help people who aren't using it in the EU.

1 more reply

marcoc2y ago

How can one create a professional looking pdf like the QOAF specification one?

GraemeMeyer2y ago

Two-column layout in Microsoft Word, large header, smaller footer, with appropriate font choices would get you basically all the way there.

jfk132y ago

HTML+CSS, converted to PDF via the Save As PDF feature in Firefox. (Or the same could be done with other browsers, but this one apparently comes from FF.)

crumpled2y ago

I looked at the PDF, and can confidently say I could typeset that in a word processor, using a stylesheet to sustain it.

That's not what they did, apparently.

The document properties call out https://cairographics.org

kenferry2y ago

Cairo’s a couple layers down from what you’re talking about. It’s the actual glyph rendering.

1 more reply

rockstarflo2y ago

What is the tradeoff there?

DamonHD2y ago

> QOA is slower than ADPCM, doesn't compress as much as MP3 and sounds worse than FLAC (duh). But I believe it fills a gap that was worth filling.

jandrese2y ago

MP3 compression is very fast on modern hardware. This may have a niche for low power devices, especially if they are battery constrained.

3 more replies

daneel_w2y ago

That in terms of quality per any bitrate it comes nowhere near ubiquitous formats like AAC or MP3 when produced with good encoders. But it's good to have (possibly) patent-free solutions available.

jychang2y ago

MP3 patents expired a long time ago, no?

1 more reply

ape42y ago

What's going to be the next Quite OK thing?

p1mrx2y ago

Quite OK Food. It tastes like sand but the shelf life is above average.

m4632y ago

Sounds like soylent. (except my direct experience with soylent leads me to think super-processed isn't that OK foodwise)

_yb2s2y ago

Sounds like freeze dried backpacking food, except at $20/meal, it's not quite OK.

mhd2y ago

The author wrote a very simple MPEG[1] decoder, so there's an obvious benchmark for making that even simpler.

I personally wouldn't mind a Quite OK Page Description Langage. Something that gets you most of PDF/PS/HPGL without all the effort. Could use the Quite OK Image Format for bitmap images. Not sure whether you'd need a Quite OK Vector Format and/or a Quite OK Font Format as prerequisites…

[1]: https://phoboslab.org/log/2019/06/pl-mpeg-single-file-librar...

dvh2y ago

Quite OK browser. It doesn't have webgl, webgpu or other fancy and easy to exploit stuff, but it renders 95% of websites and source code is easy enough to be maintained with very few people.

kragen2y ago

maybe links2 or dillo?

marmakoide2y ago

Quite OK JS Plotting Library (QOJSPL, nice, sounds like my cat walking on the keyboard). With an intuitive, documented API that doesn't require you to dig through tons of examples on sites that take ages to load. Because, no, a massive stash of non-orthogonal examples does not replace a documentation.

AKA last Tuesday morning frustration : I wanted to make interactive plots on a web page to explain math stuffs.

extua2y ago

TinyVG follows the similar goals: an alternative to SVG with a specification which trades off features for simplicity. https://tinyvg.tech/

bartwe2y ago

Hopefully a movie format

WithinReason2y ago

MPEG1 is actually quite OK

ericls2y ago

The smaple page preloads all the files before playing... Which wastes lots of bandwidth.

Aldipower2y ago

An _audio_ format which is _quite_ ok? Not sure, if I need that.

j / k navigate · click thread line to collapse

63 comments

kragen2y ago

has anyone benchmarked qoa to see roughly how many instructions per sample it needs? all i see here is that it's more than adpcm and less than mp3, but those differ by orders of magnitude

like, can you reasonably qoa-compress real-time 16ksps audio on a 16 megahertz atmega328?

so, quite plausibly

for decoding the same math gives 43 μops per sample rather than 380

i'm very interested to hear anyone else's benchmarks or calculations

g0xA52A2A2y ago

Some previous discussions.

3 months ago - https://news.ycombinator.com/item?id=35738817

6 months ago - https://news.ycombinator.com/item?id=34625573

kragen2y ago

thank you very much

these had crucial information for me

mips_r4300i2y ago

For example, instead of using two 4bit ADPCM channels for stereo, which would only be a 50% savings over uncompressed, you could probably use an average of 5 bits per sample.

anotherhue2y ago

> Would be nice to see joint stereo support

This was/is available in MP3 since forever, so seems a reasonable request.

https://wiki.hydrogenaud.io/index.php?title=Intensity_stereo

gaazoh2y ago

I like the philosophy of QOA (and other similar projects, including QOI and TinyVG), but unlike others, it seems like it's not ready to use yet, see https://github.com/phoboslab/qoa/issues/25

> I have just pushed a workaround to master. [...]

> This still introduces audible artifacts when the weights reset. It prevents the LMS from exploding, but is far from perfect :/

This, combined with the fact that that issue is still open mean that a breaking change is still to be expected.

codeflo2y ago

Pet_Ant2y ago

What is the LFE channel?

It should be spelled out explicitly, but I figured out the rest

L-Left,R-Right,C-Center,FL-Front Left,FR-FrontRight,SL-SideLeft,SR-SideRight,BL-BackLeft,BR-BackRight

---

Edit: LFE-LowFrequencyEffects... so subwoofer?

https://www.dolby.com/uploadedFiles/Assets/US/Doc/Profession...

ogurechny2y ago

LFE audio channel is different from subwoofer output.

Doxin2y ago

So it's essentially a bass shaker track?

1 more reply

entropicdrifter2y ago

LFE is an industry standard term for the subwoofer channel. It's the ".1" in "5.1","6.1","7.1" etc

ok_dad2y ago

samplatt2y ago

LFE can mean "bass shaker", but it's an industry-standard term invented by Dolby that effectively means "between 3 and 120hz", which usually means "subwoofer".

These days crossover points are very configurable. Most bass shakers are rated for use between 20hz and 200hz.

bravura2y ago

Low frequency energy, I assume. Ie bass. Your subwoofer or „bottoms“ if you have several.

MobiusHorizons2y ago

Seems to have similar design criteria as opus but I don’t see any comparison.

Turing_Machine2y ago

Other than that, looks great!

speedgoose2y ago

The website says it’s made in Hesse. No software patents to care about there.

https://en.m.wikipedia.org/wiki/Software_patents_under_the_E...

morelisp2y ago

Probably the most infamous audio format patent ever was owned by a German research institute.

1 more reply

IshKebab2y ago

1 more reply

Turing_Machine2y ago

Maybe not, but that doesn't help people who aren't using it in the EU.

1 more reply

marcoc2y ago

How can one create a professional looking pdf like the QOAF specification one?

GraemeMeyer2y ago

Two-column layout in Microsoft Word, large header, smaller footer, with appropriate font choices would get you basically all the way there.

jfk132y ago

HTML+CSS, converted to PDF via the Save As PDF feature in Firefox. (Or the same could be done with other browsers, but this one apparently comes from FF.)

crumpled2y ago

I looked at the PDF, and can confidently say I could typeset that in a word processor, using a stylesheet to sustain it.

That's not what they did, apparently.

The document properties call out https://cairographics.org

kenferry2y ago

Cairo’s a couple layers down from what you’re talking about. It’s the actual glyph rendering.

1 more reply

rockstarflo2y ago

What is the tradeoff there?

DamonHD2y ago

> QOA is slower than ADPCM, doesn't compress as much as MP3 and sounds worse than FLAC (duh). But I believe it fills a gap that was worth filling.

jandrese2y ago

MP3 compression is very fast on modern hardware. This may have a niche for low power devices, especially if they are battery constrained.

3 more replies

daneel_w2y ago

That in terms of quality per any bitrate it comes nowhere near ubiquitous formats like AAC or MP3 when produced with good encoders. But it's good to have (possibly) patent-free solutions available.

jychang2y ago

MP3 patents expired a long time ago, no?

1 more reply

ape42y ago

What's going to be the next Quite OK thing?

p1mrx2y ago

Quite OK Food. It tastes like sand but the shelf life is above average.

m4632y ago

Sounds like soylent. (except my direct experience with soylent leads me to think super-processed isn't that OK foodwise)

_yb2s2y ago

Sounds like freeze dried backpacking food, except at $20/meal, it's not quite OK.

mhd2y ago

The author wrote a very simple MPEG[1] decoder, so there's an obvious benchmark for making that even simpler.

[1]: https://phoboslab.org/log/2019/06/pl-mpeg-single-file-librar...

dvh2y ago

Quite OK browser. It doesn't have webgl, webgpu or other fancy and easy to exploit stuff, but it renders 95% of websites and source code is easy enough to be maintained with very few people.

kragen2y ago

maybe links2 or dillo?

marmakoide2y ago

AKA last Tuesday morning frustration : I wanted to make interactive plots on a web page to explain math stuffs.

extua2y ago

TinyVG follows the similar goals: an alternative to SVG with a specification which trades off features for simplicity. https://tinyvg.tech/

bartwe2y ago

Hopefully a movie format

WithinReason2y ago

MPEG1 is actually quite OK

ericls2y ago

The smaple page preloads all the files before playing... Which wastes lots of bandwidth.

Aldipower2y ago

An _audio_ format which is _quite_ ok? Not sure, if I need that.

j / k navigate · click thread line to collapse