Commits as Snapshots
Git commits as snapshots are immutable objects, not saved patches. A commit names the complete tree built from the index, records its parent, author and committer data, and message, then hashes that serialized payload. Change only the parent and the files can stay identical while the commit ID changes.
Think of the tree as a photograph and the commit as the dated card taped to its back. Two cards may carry the same photograph, but if one says "previous card: 309aeee" and the other does not, the cards are not interchangeable. Git hashes the whole card, not just the picture.
How Git commits as snapshots work
The index is the source of the snapshot. Each staged path points to a blob object, and each blob ID comes from the file's exact bytes. The tree object records the full set of path-to-blob entries. A commit then records the tree ID, an optional parent ID, author and committer lines, and the message.
That gives Git three content-addressed object types in this lesson:
- A blob stores file bytes. Its filename is not part of its identity. Two files with identical bytes can use the same blob.
- A tree stores names, modes, and object IDs. Change one file and the next tree can still point to every unchanged blob from the previous tree.
- A commit stores a tree ID and history metadata. Its parent field turns an otherwise identical project snapshot into a new point in a specific history.
The hash recipe is exact. Git prefixes an object's body with its type, a space, the byte length, and a null byte. It then hashes header plus body. In the scoped SHA-1 format used here, the empty tree is always 4b825dc642cb6eb9a060e54bf8d69288fbee4904. That value is not a friendly label invented for the diagram; it is the real Git object ID.
Object IDs are identities, not storage slots waiting to be edited. If Git sees the same bytes again, it gets the same ID and can reuse the object already in the database. If one byte changes, it writes a different object. The old one stays valid because nothing rewrites it.
This model explains several everyday Git surprises. git diff computes changes between two snapshots; it is not reading a patch stored inside either commit. Rebase, cherry-pick, and git commit --amend can leave the checked-out files unchanged while producing a new commit ID because ancestry or metadata changed.
Commits as snapshots step by step
The default run uses app.py and README.md, a linear main, and fixed identity and time fields. Fixing the metadata makes the comparison honest: when the final commit ID changes, you can point to the parent field instead of blaming a clock.
- Edit
app.pytoprint("one"). This changes working bytes only. There is no blob yet. - Add
app.py. Git hashes those 12 bytes into blob51ee69d..., stores the blob, and writesapp.py -> 51ee69d...into the index. - Edit
README.mdtohello, then add it. The bytes become blobb6fc4c6..., and the index now has two paths. - Commit with message
base. Git serializes both index entries into treed5ba943.... The first commit has no parent, so its payload names that tree, the fixed identity lines, andbase. The resulting commit begins45ecc20...;mainmoves there only after the object exists. - Edit
app.pytoprint("two")and add it. That creates a new blob forapp.py. Nothing touchedREADME.md, so its index entry still points tob6fc4c6.... - Commit with message
update. The new treeff8ad65...contains the newapp.pyblob and the oldREADME.mdblob. The commit names45ecc20...as parent and hashes to309aeee.... - Commit again with
--allow-emptyand the same messageupdate. The index has not changed, so Git reuses treeff8ad65.... The message, identity, and time also match the previous commit. One field differs: the parent is now309aeee.... That different payload hashes toccf8be2..., andmainadvances once more.
The sixth step is structural sharing in plain sight. Git does not copy the unchanged README bytes into a fresh container. The new tree simply points to the same blob. The seventh is the sharper result: the project snapshot does not change at all, yet history does.
Why identical snapshots have different commit IDs
A commit ID authenticates more than the files you can check out. It authenticates the commit payload, and that payload includes ancestry. C2 and C3 in the default run both point to tree ff8ad65..., but C3 names C2 as its parent while C2 names C1. Their histories are different, so their IDs must be different.
Remove the parent line and the model breaks in a specific way. Suppose two repositories have the same tree, fixed identity and time, and the same message, but reached that state through different previous commits. Without the parent field, both would serialize the same commit bytes and receive the same ID. The ID would no longer prove which prior history the commit follows.
This is why a commit is not merely a tree with a message. The tree answers, "Which snapshot?" The parent answers, "After which history?" The hash seals both answers together.
This also explains --allow-empty. The flag does not manufacture a duplicate tree. It bypasses the ordinary equality guard and permits a new commit payload that points to the existing tree. The new parent field is enough to give that payload a new identity. It is useful for markers, automation triggers, and recording an event with no file change, but it is not the normal way to save work.
Commit Snapshot Edge Cases
- Nothing staged. A normal commit compares the tree built from the index with the tree at
HEAD. If they match, it stops before assembling a commit payload. A dirty working file does not matter untilgit addupdates the index. - Allow-empty commit. The same tree is accepted, but the current
HEADstill becomes the new parent. A new commit ID follows even when message and fixed metadata also match. - Missing working file. For this path-specific
add, a missing source creates no blob and no index entry. Real Git also models deletions, but deletion staging is outside this two-file engine. - Identical file bytes. Adding the same bytes again produces the same blob ID. The database does not need a second blob card.
- Short ID collision. The player displays the shortest unique prefix with a minimum of seven characters. Git object identity is still the full hash; a prefix may need to grow when another object begins with the same characters.
- Hash format. This lesson uses Git's SHA-1 object format. Git repositories can also use SHA-256, so do not write tools that assume every object ID is 40 hexadecimal characters.
Common Mistakes About Commit Snapshots
"A commit stores changed lines." Diffs are useful views computed between snapshots. A commit points to one complete tree. Unchanged paths remain because the tree still names their blobs.
"Changing one file copies the whole repository." The new tree is a fresh path map, but unchanged entries reuse existing object IDs. Git's object graph shares structure aggressively.
"The commit ID is a hash of the files." It is a hash of the serialized commit body, which contains a tree ID, parent, identity lines, and message. The files affect it through the tree, but they are not the whole payload.
"The parent is just a navigation hint." It is inside the bytes being hashed. Delete or change it and you have a different commit object, even if checkout would produce the same files.
"--allow-empty creates an empty tree." It normally reuses the current staged tree. "Empty" means no file-tree change relative to the parent, not a repository with no files.
"Git recomputes IDs from whatever is on disk now." The working tree is outside the object database. IDs come from stored object bytes and remain stable while your editor keeps changing files.
A note on the scoped model
The animation uses two flat files, mode 100644, linear history, fixed author and committer data, and deterministic SHA-1 object serialization. It omits directories, executable and symbolic-link modes, deletions, merge parents, signatures, hooks, branch/ref storage details, index metadata, SHA-256 repository setup, packs, deltas, reflogs, and extensions. Those features add more object and reference machinery; they do not change the central rule that a commit hashes a tree, its ancestry, and its metadata.
For the full object formats and command behavior, see the Git object internals chapter, git commit-tree, and Git's hash function transition plan.
Keep exploring
A few lessons that share the same technique or lead into the next idea.
Found a mistake or something unclear? Leave feedback.