Skip to content

Storage

The disk holds two worlds: a small FAT firmware partition, and the HFS+ data partition where the music, the databases and the artwork live. Chikuma drives both, and it drives them in the arrangement the device uses - which, as the kernel page explains, is not the obvious one. The code is in chikuma/drivers/fs/ (and the ATA driver in chikuma/drivers/hw/).

The layers

From the disk up:

  • ATA. Reads are reproduced for both 28-bit and 48-bit addressing, with the DMA read path modelled. Writes use the firmware's own programmed-I/O fallback forms; the DMA write command is recovered but not yet reproduced. The 28-bit guard refuses an over-range write rather than masking the address, because a masked write would land on the wrong sectors. See the device models for the disk hardware underneath.
  • The partition layer. Chikuma walks the Apple Partition Map to find the data region. The data partition is itself nested: it carries its own partition map inside, and a flat HFS+ volume with no inner map will not register - which matches how a real synced device is laid out.
  • FAT. FAT16 read and write are reproduced for the firmware resource partition. For a blank or foreign disk the firmware creates a FAT32 volume, and Chikuma reproduces that self-formatter; it never writes HFS.
  • HFS+. The read side, in-place writes, growing a file (allocation bitmap plus the extents-overflow tree), create, delete, rename, and extended attributes are all reproduced, along with a journal writer and replayer that is Chikuma's own (see below).

4K logical sectors: the trap that fails silently

The iPod disk presents 4096-byte logical blocks, not 512. The filesystem requests in 4096-byte units and the block layer scales to the device's own sector size. This matters because the single read the firmware issues to bring up the volume is one 4096-byte transfer at the start of the data partition, and the HFS+ volume header (which sits 1024 bytes in) falls inside that one block. An image built on a 512-byte assumption puts the partition and the header at the wrong byte offsets and the volume fails to mount with no obvious cause. Chikuma's images are built with the 4K geometry, and there are in fact three sector sizes in play at once - the ATA transfer unit, the partition map's unit, and the FAT parameter block's own - which is a documented source of confusion.

The filesystem posts to a queue and blocks

This is the load-bearing architectural fact, and it is worth stating twice because it is so easy to get wrong. The filesystem does not call the disk. It builds a request, pushes it onto a work list, and blocks. A separate disk task issues the command later, on its own stack, and the completion interrupt wakes the caller.

Chikuma reproduces this seam exactly, because it is how the device is arranged and because structure follows the firmware. The consequence for anyone tracing the system: a stack trace taken at the moment a disk command is issued can never name the filesystem code that wanted the data, because that code is in another task, already blocked. The full request chain was recovered by running it and watching the queue, not by reading a call graph - the two ends are in different tasks by construction. The read and write drivers for FAT and HFS+ are told apart the same way, by behaviour, because the seam between them is built at mount rather than living as a fixed table in the image.

HFS+ in detail

HFS+ is a published on-disk format (Apple's Technical Note TN1150, and Apple's own open-source HFS sources), so Chikuma treats the format knowledge as a standard and keeps only one thing recovered: which operations the firmware performs, in what order, with what arguments. The volume header fields, the four special forks, the catalog and extents B-trees, the fork layout and the big-endian on-disk encoding are all read the way the standard defines them, and the recovered-from-firmware values were cross-checked against Apple's published format headers and agreed.

  • The catalog reader walks the catalog B-tree: header node, the node size, records stored back to front from each node's end, the parent-id-plus-name keys, the folder/file record types, and a file's data fork. All big-endian.
  • The B-tree machinery is tree-agnostic and shared by the catalog, extents-overflow and attributes trees. Node splitting is done by record count and deliberately omits borrow-from-a-neighbour - the "simpler always-correct half", which matches Apple's own code, which also does not rotate. A synthetic split test that forces new root nodes found a real bug (a double record-count increment on the new-root path) that 500-file integration runs never surfaced.
  • Dates. The catalog carries exactly three date fields Chikuma writes - create, content-modified and backup - stamped in UTC, using the standard 1904-to-1970 epoch offset. Two further date fields exist in the format but nothing measured fills them, so reproducing five would be inventing two.

The journal is Chikuma's, not the firmware's

The firmware writes a journaled HFS+ volume but never maintains the journal - measured directly (the journal's magic never appears in the image, and first-boot writes never land in the journal region on the emulated machine). So the journal writer and replayer in Chikuma are declared OURS. They exist for a practical reason: current macOS silently discards a pending journal on a volume with allocation blocks larger than 4K (which the iPod's geometry is), so Chikuma replays its own journal at mount - validate-first, all-or-nothing - and the write path with journalling turned off does exactly what the firmware does.

Row order comes from the database, not from us

The order rows appear in - Songs, Albums, Artists, Genres, Composers - is the database's order, not one Chikuma computes. iTunes writes already-sorted index lists into the iTunesDB, and the device builds its lists from those and then filters and groups without re-sorting. Chikuma does the same: it walks the file's own artist index and takes each album's first appearance, which is what fixed the Cover Flow and Albums ordering without decoding a single sort-descriptor field. Where the device does compare strings (a Songs list, which has no distinct-value index), it uses a Unicode-weight collation - punctuation before symbols before digits before letters, with case and accent as secondary weights - which Chikuma reproduces from a weight table read off a running device. This is all written up in reports/library_ordering.md.

Tests, and what is open

Tests run against the real volume image and real iTunesDB databases, never against fixtures written to match the parser. For writes, the oracle is macOS's own fsck_hfs, run on a temporary copy of the image, and negative controls (deliberately corrupted volumes) are run so that a gate which never fails is not trusted. Several real-world bugs on the delete and split paths were found this way.

Honestly open: the DMA write direction; the mount gate itself (what turns a found partition into a mounted volume is tied to device discovery and is not fully recovered); splitting a fully-populated catalog leaf on a create is refused rather than performed on some paths; and the write path has not been exercised on real hardware. The live account is in Status, and the per-function state is reports/fs_worklist.md.