Thirty Days of Grok Heavy Build
Running Grok 4.5 Build under SuperGrok Heavy, inside a Claude-native harness I started in December 2025 · July 8 – August 8, 2026
I spent the first ~30 days of a SuperGrok Heavy subscription running Grok 4.5 Build — the agentic CLI — inside a harness I built first for Claude Opus, then integrated Codex, and later other model routes for volume versus judgment.
This is not a marketing feature list. It is what the on-disk session store shows when every Heavy chat history in that window is inventoried and pushed through a redaction-aware mine.
The month, in one table
| Window | 2026-07-08 → 2026-08-08 |
| Sessions | 7,342 (~2.85 GB of session transcripts) |
| Tokens (Heavy) | ~7.37B (prompt-heavy; local daily meter, 27 token-tracked days Jul 13–Aug 8 — slight undercount possible early in the window) |
| Tokens (all Grok accounts I ran that month) | ~9.45B |
| Subscription vs API counterfactual | $99 specialty 3-month Heavy promo vs ~$15,300 at Grok 4.5 list rates ($2/M input · $6/M output; reasoning counted as output; no long-context, tools, or cache adjustments) — about 155× |
| Interactive sessions | 577 |
| Headless worker sessions | 4,478 |
| Worktree sessions | 1,502 |
| Dominant model | Grok 4.5 |
Headline: Heavy Build was not a chat toy in my harness. It was a fleet co-processor — thousands of headless workers, hundreds of interactive sessions, and a habit of treating “done” as something I re-check on disk.
Names
| Name | Meaning | Not |
|---|---|---|
| Grok Build (CLI) | Terminal agent with tools, sessions, and detached workers | Web or mobile “Build Mode” |
| SuperGrok Heavy | The subscription I was on | Metered API keys |
| 4.5 | Model family for almost all of this work | Older Grok chat models |
| Harness | My local agent OS — projects, job routing, durable notes, live meters | The model weights alone |
A foreign harness, on purpose
Most demos put a model in its own living room. I did the opposite: a system designed around Claude and Codex, with Grok Heavy as a high-capacity, tool-capable family and multi-account routing. Cheap volume went to detached workers; judgment and single-writer artifacts stayed with me.
If Build only works inside a pure xAI demo loop, it is a product. If it thrives inside a foreign, fail-closed operator stack, it is infrastructure.
Here, it proved to be infrastructure.
Scale is a catalog
I did not ask a model how many sessions I had. A script walked the Heavy session store, classified working directories, and wrote a catalog.
Most sessions were not long chats with me. The median was a bead: a detached, prompt-filed worker writing into an isolated output directory, then harvested by me in a controlling session that does not trust a “complete” flag without re-reading the files. That is the product story — Build as bacterial labor, not only as a pair programmer.
Interactive sessions were the steering surface: automations, meters, design, games and visual critics, capacity routing, adversarial verifiers. One peak day (2026-07-28) was almost entirely that kind of work — a studio day and a control-plane day sharing one account envelope.
Multi-track month
Capacity retrospectives fail when they collapse parallel programs into one story. This month ran simultaneous tracks: fleet orchestration; content and education; games and visual QA; systems truth (usage meters, multi-model accounts, host load, routing).
Keyword co-occurrence across takeaways (not exclusive labels) put verification and orchestration first, with content, capacity, games, and infrastructure close behind. A single session can hit several themes.
What Build was good at here
Long-horizon work with tools. Multi-hour arcs: research → filesystem → reports → notes. Not just single-turn Q&A.
Headless volume without losing the plot. Workers for scouts, competitive profiles, fix passes, reviews — with isolation, registration, and a re-check of the files before I accept them.
Adversarial self-check. Separate Build sessions tasked to refute “objective complete.” Optimism is cheap; refutation is culture.
Capacity as a portfolio. Heavy is a metered weekly envelope you aim, not a bottomless key. Account routing and burn honesty mattered as much as prompts.
Scars worth shipping
- “Complete” is a claim. Disk re-check is the gate.
- Prose is not state. Counts live in catalogs; status prose is a pointer.
- Parallel authors fracture voice. Coupled narrative stays sequential.
- Packaging ≠ publish. A human stamp is separate from a finished draft.
- One account blocked is not a fleet halt. Re-route; re-observe live meters.
- Host surge is real. Wide fan-out under heavy load manufactures flakiness.
- Redaction before fame. Full-session mines touch private work; public layers only consume what is tagged safe.
How this review was made
Honesty about method is part of the story:
- Deterministic catalog of every Heavy chat history in the window
- Deterministic digests (title, asks, theme hits, sensitivity heuristic) for all of them
- Fifty size-balanced shards → cluster → chronology and capability integrate
- One sequential write of this piece
Why digests before generative judgment: feeding multi-gigabyte raw JSONL into dozens of concurrent models under a loaded host is a reliability failure dressed as ambition. The pyramid still exists; the base is evidence-compressed.
Roughly half the takeaways stayed out of the public chain under redaction heuristics (including known false positives). Better blocklists can refine later without re-mining raw text into drafts.
What I would tell a new Grok Build user
- Put Build in a real harness — output contracts, hard checks before you call work done, notebooks.
- Prefer detached workers for volume, interactive for judgment.
- Inventory sessions before storytelling.
- Treat capacity as a portfolio, not a single chat.
- Ship practices in the same motion as demos.
- Never auto-publish from agent enthusiasm.
Closing
Thirty days in, Grok 4.5 Build under Heavy is not a sidebar experiment in my operator OS. It is a primary pair of hands for fleets, a critic for its own optimism, and a co-author of durable practice.
The piece is the month: thousands of sessions, one redaction-aware pyramid, and a public story that refuses to invent what the disk does not show.
Get new posts by email — first
The newsletter is in the works — join the waitlist and be first to know when it launches. Everything here stays free to read.
The Solo Stack is written by Matt — building products solo with AI, on his own infrastructure. If a claim isn’t backed by experience or a measurement, it doesn’t ship.
Not sending yet: joining stores your address on the waitlist. One confirmation email at launch — nothing sends unless you confirm.