Measure pressure
Codex compares active token usage with the auto-compact limit and the model's usable context window.
When a thread grows too large, Codex swaps the model's working transcript for a smaller checkpoint. The saved rollout remains intact. What changes is the history sent into the next inference request.
Compaction is a checkpoint operation. Codex detects pressure, picks one of three compaction routes, replaces the active model history, and records enough state to resume from that new boundary.
The details differ by provider and feature flags. The outer shape does not.
Codex compares active token usage with the auto-compact limit and the model's usable context window.
A token-budget reset wins first. Otherwise provider capability selects remote or local compaction.
The new history contains a summary or opaque checkpoint plus selected messages and fresh context.
Codex advances the context-window ID, persists replacement_history, recalculates usage, then resumes work.
Pick a tab. The current source checks these routes in this order.
Codex copies the current prompt input, attaches known tool calls, and appends a CompactionTrigger item. The provider must return exactly one ResponseItem::Compaction.
Codex appends its built-in checkpoint prompt, runs a normal model response, and treats the last assistant message from that compaction turn as the summary.
This experimental route skips model and server summarization. It rebuilds initial context from current settings and world state, then installs that as the replacement history.
Token pressure is the common case, but it is not the only reason Codex compacts.
At the start of a user turn, Codex checks the existing history. If it is already at the threshold, compaction runs before the new input is recorded.
After a model response, Codex compacts only when more work remains and the limit was reached or the model requested a new context window.
A changed compaction compatibility hash triggers compaction. Moving to a smaller context-window model can also trigger it when current usage no longer fits.
/compactThe TUI exposes a direct command. It creates a standalone compaction turn and uses the same provider and feature routing as automatic compaction.
auto = min(configured limit, 90% of context)
body_after_prefix scope, Codex subtracts the current window's initial prefix before comparing usage with the configured auto-compact budget.Compaction preserves task continuity, not a verbatim transcript. Each route has different retention rules.
| Input | Remote v2 | Local summary | Token-budget reset |
|---|---|---|---|
| User messages | Selected and retained, newest-first budget when truncating | Recent messages retained within 20k tokens | Not carried by the compaction operation |
| Assistant and tool history | Compressed into the opaque provider item | Compressed into the handoff summary | Not carried |
| Developer messages | Optional for client-authored items | Reinjected as current initial context | Optional for client-authored items |
| Current settings and world state | Reinjected at the appropriate boundary | Reinjected at the appropriate boundary | Installed as full fresh context |
| Images | Budgeted retention behind a feature flag | Not copied into the text summary history | Not carried by the reset itself |
The old rollout records stay on disk. Codex appends a CompactedItem that names the new model-visible starting point.
The checkpoint stores replacement history, window lineage, the remote response ID when there is one, retained review context, and a token-usage snapshot.
For paginated rollouts, the reverse scanner can stop once it finds a usable compaction plus completed turn context. It does not need to replay the thread from its first line.
{
"type": "compacted",
"replacement_history": [
// selected messages + summary/checkpoint
],
"window_number": 4,
"previous_window_id": "…",
"window_id": "…",
"compaction_response_id": "resp_…",
"latest_token_usage_record": { /* … */ }
}
The source has guardrails, but compression still trades detail for room.
Codex uses server usage when available and local estimates elsewhere. Exact trigger timing can differ from what a user infers from visible text.
Tool outputs, intermediate reasoning, and older messages may survive only through compressed task-relevant information.
The local route warns that long threads and multiple compactions can reduce accuracy. A focused new thread remains the cleanest reset.