Audio and video¶
The most surprising thing about media on this device is how little of it is software. Both audio and video are decoded by dedicated hardware blocks on the chip. That single fact shapes the whole subsystem and it is the clean-room point worth stating first: there is no Apple codec in the image to recover. The silicon decodes. What Chikuma reproduces is the contract with that hardware - how it is fed, how it signals, and what it hands back - not a decoder.
Audio: feed the block, wait, collect¶
The audio decoder is a bus-mastering hardware engine. The firmware's job, and therefore Chikuma's, is:
- Configure the engine for the track from the file's own sample description.
- Hand it a linked list of descriptors in main memory, each pointing at one compressed packet.
- Wait on the completion interrupt.
- Collect the decoded 16-bit PCM from the buffers the engine announces.
Compressed audio goes in over the bus, PCM comes out. The block is a generic transform engine with no
notion of which codec it is running - the codec lives entirely in how the elementary stream is framed
(an MP3 frame is handed over whole; an AAC payload has its container framing stripped first). This was
confirmed by comparing the engine's entire configuration between an AAC track and an MP3 podcast: only a
handful of numeric parameters differed, and nothing named a format. So to make the emulator produce
sound it has to be the decoder (it decodes host-side), and to run on real hardware Chikuma has to
program the block. Either way there is nothing of Apple's decoding logic to describe, only the
interface, which is specified in reports/sm1_spec.md and reproduced in chikuma/drivers/audio/.
The feeder blocks, it does not spin. It hands over a packet and waits on a semaphore that the completion interrupt signals; a non-empty decoded result keeps it streaming, an empty one ends the track. Blocking rather than spinning is what lets the window manager keep painting while a track plays - Now Playing updates its title, position, progress and cover with playback running underneath.
A field nothing writes is the hardware's¶
Two audio crashes taught the same lesson, and it now governs the whole device-model effort: if the firmware reads a value it never wrote, that value is the hardware's to publish, and the model owes it as an input.
- The engine's output side has a small ring of buffers, and the firmware wraps its index against a ring size it never writes - the hardware provides it. Reading zero there meant the index never wrapped, so the buffer address walked off the end of memory and the firmware faulted. This is what a "stall at 200 buffers" actually was: a crash of our own making, fixed by having the model publish the count.
- A pair of position values describe a window over the current output, not a running total. An early model treated one of them as a cumulative counter, so a derived pointer drifted out of memory and the firmware took a data abort at exactly end-of-track, several minutes in. The fix was to model the field as what it is.
Video: a separate block, fed slices¶
Video has nothing to do with the audio engine - measured directly, by decoding a film and a song and confirming the audio engine's traffic is byte-for-byte identical between them. Video is its own on-chip hardware H.264 decoder, and like the audio engine it is hardware, so again there is no codec to recover, only the handshake, the buffer layout and the output contract.
It is not fed raw H.264 network units. It takes a per-slice descriptor table: each entry carries the
slice's already-stripped body plus the header fields the driver parsed, and the decoder reconstructs
from those. Reproduced in chikuma/drivers/video/, it now:
- Decodes cleanly, including a two-layer temporal GOP. A reconstruction drift was traced to how one reference-flag field is derived; taking it from the frame's position in the group made every frame decode with zero errors, and the rebuilt units self-check against the descriptor.
- Starts on the right edge. The run control is a bit that stays set across frames; the hardware acts on the rising edge and clears the bit itself on completion. The model has to clear it on completion too, or the next frame never presents a rising edge and the stream desynchronises. This was the timing bug that only the instant-interrupt QEMU machine exposed.
- Reaches the panel, full screen. The decoded frame is composited on the one display-pipe window that scales rather than crops, through a hardware polyphase scaler, with fit-to-screen versus full-screen taken from the Video setting.
Protected films stay black, deliberately. A FairPlay-encrypted film is fed ciphertext, the unit self-check rejects it, and the picture is correctly black. The reports stop at structure; no decryptor is built, and that is a project rule, not a gap to be closed.
The 3D is software¶
Cover Flow and the tilted album art are drawn by the CPU, not a GPU. The image's OpenGL ES is Vincent, a software rasteriser, which the device runs with a just-in-time backend that generates code at run time. Chikuma uses Vincent from upstream at the version the image advertises, and recovers only the one vendor extension Apple grew onto it. So the tilted, reflected covers are the processor's work on both the original and Chikuma, and they are compared frame to frame against the original on the same flick.
Volume, from the wheel to the codec¶
The volume path is recovered end to end rather than approximated, and it is a good example of a chain that reaches all the way from an input event to a chip register:
- A wheel notch moves the level by two, clamped by a ceiling.
- The level lives on a 256-step internal scale, even though far fewer distinct codes are audible at the far end.
- There is a region ceiling that is the board's, not a setting: a unit sold into a market with a volume cap tops out lower, and the user's own volume limit can never climb past it because the step-up clamps against it.
- The level maps through a fixed curve to the audio codec's own output-gain registers (a Wolfson-style
part's
LOUT1/ROUT1), written over the two-wire (I2C) bus with the codec's update-latch bit so both channels apply together. The reproduction matches a measured bus capture exactly, including why the 256 input levels collapse onto far fewer distinct codes. - Pause is one byte. The device's pause clears a single flag and leaves the transmitter configured and the cursor in place, so playback resumes without reprogramming - there is no DMA teardown at all. Stop, by contrast, tears the transmitter down and rewinds. Chikuma's pause reproduces this subset and says so.
One calibration byte in the curve has no source in the image and is named as zero in the code, which is the kind of stand-in the project marks rather than guesses.
The headphone jack¶
Detection and the in-line remote are largely reproduced (chikuma/drivers/hw/mikey.zig). The headset
remote is a small controller on the I2C bus; presence is read from a GPIO line, but only after a one-shot
task has published the plug object, which is why a bare emulated device reads "unplugged" until that runs.
The remote's play/pause and transport buttons were confirmed on real hardware. The interrupt is held
until the firmware's own read clears it, which the emulator model reproduces and a gate asserts; the
analog wire protocol below the decoded button-and-channel level is outside the recovery boundary. Details
are in reports/headphone_jack.md.
What is open¶
The audio codec's per-track parameter fields are only partly understood (two modes are confirmed by trace, the rest identified only by structure, for want of media to exercise them); the ported feeder's configuration front-end and give-up timeout are not yet wired; and the protected-media transform path is documented but deliberately never built. On the video side, the block plays the unprotected film and the open questions are narrow (one quantisation field, and confirmation on streams that vary it). The live account is in Status.