mirror of
https://github.com/Rockbox/rockbox.git
synced 2026-10-10 08:03:04 -04:00
inverse_channel_transform() ran its general N-channel matrix loop for every sample of a stereo stream, where the matrix is exactly +-1.0. That loop was a quarter to a third of the whole decode, and GCC 9.5.0 compiles it worse than 4.9.4 did, which made the codec 6-9% slower on ARM7TDMI after the toolchain update. Handle a group of two channels separately: add and subtract when the matrix is +-1.0, and a plain four multiply loop otherwise. More than two channels still use the general loop. This reverses the regression from the GCC 9.5.0 update and goes well past it. The loop the newer compiler handled badly is no longer used for stereo, so the two compilers now give the same speed to within 1%, about 25% faster than the codec was with GCC 4.9.4 (estimated with perfsim, e200v1, wmapro_141k: 25.81 MHz with 4.9.4 before this change, 19.5 MHz with either compiler after it). Output is bit-identical: whole-file PCM hashes match before and after for five stereo files at 55-271 kbps, built with GCC 9.5.0 and with 4.9.4, and also with the multiply path forced on. Measured with test_codec, wmapro_141k.wma, MHz for real time: Sansa e200v1 27.99 -> 19.70 Sansa Clip+ 21.78 -> 15.80 Estimated with perfsim for the other files (e200v1 / Clip+): wmapro_55k 25.21 -> 17.06 / 20.02 -> 13.71 wmapro_80k 26.17 -> 18.01 / 20.75 -> 14.44 wmapro_173k 28.52 -> 20.29 / 22.55 -> 16.25 wmapro_271k 30.85 -> 22.52 / 24.34 -> 18.04 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| arm_support | ||
| fixedpoint | ||
| libsetjmp | ||
| microtar | ||
| mipsunwinder | ||
| rbcodec | ||
| skin_parser | ||
| tlsf | ||
| unwarminder | ||
| utf8proc | ||