A WMA Professional stream of 16 bits per sample played about 48 dB
too quiet and with a noise floor near -70 dBFS.
ffmpeg scales the transform's output by the stream's sample size.
That was dropped when the decoder was converted to fixed point
(d884af2b99, 16284ae8ae), and the output is passed to the DSP as if
every stream had 24 bits. A stream of fewer bits has a lower
quantization step to match, so it came out low by the difference,
and it used the bottom of the integer quantization table, where the
factors have only a few significant bits.
Decode a 16 or 20 bit stream at the level of a 24 bit stream: use
that stream's quantization step, and scale each band's factor by
the ratio that is left. 24 bit streams are not affected.
Checked with perfsim (Sansa e200v1 build) against ffmpeg's decode
of a 16 bit, 192 kbps stereo file, the only such file to hand:
level SNR vs ffmpeg noise
before -48 dB 37 dB -70 dBFS
after correct 85 dB -117 dBFS
A 24 bit file's output is byte-identical before and after. The
20 bit case follows the same rule but is untested.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
inverse_channel_transform() ran its general N-channel matrix loop
for every sample of a stereo stream, where the matrix is exactly
+-1.0. That loop was a quarter to a third of the whole decode, and
GCC 9.5.0 compiles it worse than 4.9.4 did, which made the codec
6-9% slower on ARM7TDMI after the toolchain update.
Handle a group of two channels separately: add and subtract when
the matrix is +-1.0, and a plain four multiply loop otherwise. More
than two channels still use the general loop.
This reverses the regression from the GCC 9.5.0 update and goes
well past it. The loop the newer compiler handled badly is no longer
used for stereo, so the two compilers now give the same speed to
within 1%, about 25% faster than the codec was with GCC 4.9.4
(estimated with perfsim, e200v1, wmapro_141k: 25.81 MHz with 4.9.4
before this change, 19.5 MHz with either compiler after it).
Output is bit-identical: whole-file PCM hashes match before and
after for five stereo files at 55-271 kbps, built with GCC 9.5.0
and with 4.9.4, and also with the multiply path forced on.
Measured with test_codec, wmapro_141k.wma, MHz for real time:
Sansa e200v1 27.99 -> 19.70
Sansa Clip+ 21.78 -> 15.80
Estimated with perfsim for the other files (e200v1 / Clip+):
wmapro_55k 25.21 -> 17.06 / 20.02 -> 13.71
wmapro_80k 26.17 -> 18.01 / 20.75 -> 14.44
wmapro_173k 28.52 -> 20.29 / 22.55 -> 16.25
wmapro_271k 30.85 -> 22.52 / 24.34 -> 18.04
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No known samples are fixed by this problem, but I haven't tested many.
Backport of ffmpeg revision 26388.
Change-Id: Ife9654b7477a432834e3cab2cb43d16da071445a