The planes the de-tiling produces are already the YUV420P swscale wants, and our
sceMpeg HLE converts the same frames the same way, so the pixel formats and the
studio-range setup come straight from MediaEngine::getSwsFormat. It is 3-4x
quicker than going a pixel at a time: 0.37-0.44ms a frame becomes 0.09-0.12ms,
which is the whole reason sceMpegBaseCscAvc was at the top of a profile.
Chroma is upsampled with SWS_POINT rather than the HLE's SWS_BILINEAR, since
replicating is what the scalar path does and, being a fixed-function block,
almost certainly what the hardware does.
It is not bit-identical - swscale rounds its own way. TestMpegCsc measures the
gap per channel rather than per byte, so the number means something for a packed
16-bit pixel: worst 1 step of 31 for 5650 and 5551, 2 of 15 for 4444, 3 of 255
for 8888, with means around a fifth of a step. The scalar path stays as what the
longhand reference is checked against, and takes anything swscale won't - an odd
range origin, or a build without ffmpeg.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>