254 gas
Our Cosmos Hub validator stopped signing at 02:00 KST and did not sign again until 11:29. It was never jailed and never slashed, but it came within 37% of the line. The cause was a single number: 254 gas.
What broke
We ran the node from an image we built ourselves, and that build came out linux/arm64. The official Cosmos Hub release is published for linux/amd64 only, and that is what every other validator on the network runs.
For three hours this made no difference. Then block 32,454,059 arrived carrying a transaction that our build metered differently:
panic recovered in runTx
err="out of gas in location: WritePerByte;
gasWanted: 259076, gasUsed: 259330: out of gas"
Two hundred and fifty-four gas over the limit. On every other node the transaction fit and succeeded. On ours it ran out and failed, and a failed transaction changes no state. From that block onward our application state was a different thing from the chain's.
CONSENSUS FAILURE!!! wrong Block.Header.AppHash Expected D018907A... ← what our node computed got 56826755... ← what the network agreed on
Consensus asks for byte-identical results. Running "the same version" is not the same thing as running the same binary.
Why it took six hours to notice, and three more to diagnose
Once the state diverged, our node treated every correct block it received as invalid — and disconnected the peer that sent it.
Stopping peer for error err="reactor validation error: wrong Block.Header.AppHash"
Ninety-nine percent of the log was peers connecting and dropping. The symptom read as the network is rejecting us. We were rejecting the network.
The node also reported catching_up: false the whole time. The block-sync reactor had already exited, so as far as it knew there was nothing to catch up on. It was four thousand blocks behind and frozen.
One command would have shown the panic immediately:
tail -400 <node log> 2>&1 | grep -v 'module=p2p' | tail -40
That is now the first thing we run on any node incident.
Two things we got wrong on the way
We blamed memory. The host showed 31 of 32 GB used and heavy swapping, so we capped the memory limits on the nodes. Three healthy nodes restarted for nothing. Measured properly, the nodes were using 15.6 of 20 GiB together and the Cosmos node 3.1 of 8. The host figure was page cache from blockchain disk I/O — something we had already written down elsewhere and did not think to reread.
Then we blamed peers. We replaced a working peer list with twenty addresses harvested from public RPC nodes. Many were sentries that drop unknown connections on sight, and because persistent peers are retried forever, they occupied the dial slots and starved normal peer discovery. The original list was never the problem.
Both detours came from reading symptoms instead of reading the log.
What actually fixed it
State sync stalled three times — once mid-transfer, twice on light-client verification. A 43 GB snapshot download was prepared and never needed.
What worked was copying the data directory from another machine already running the official amd64 image, after stopping it so the database was consistent. Copying a running node's database gives you a torn snapshot and reproduces exactly the mismatch you are trying to escape.
The first copy still failed. We swapped the data but left the arm64 binary in place, and it replayed 233 blocks and diverged again at the same transaction. The binary has to change first, then the data.
Swapping to the official image needed two overrides — its entrypoint is ["gaiad", "start"] where ours was ["gaiad"], and it runs as a non-root user against a root-owned data directory. After that, uname -m reporting x86_64 was the check that mattered.
What changed
Mainnet nodes run official release images. No self-built binaries. Where the host architecture differs from the release, we run the official build under emulation — slower, and worth it. A slow node signs late; a wrong node cannot sign at all. Our Celestia nodes were never affected because they had used the official image from the start.
Our monitoring alerted 65 minutes in, then went quiet for six hours. Every rule fired as written — the rules were wrong. Alerting once per condition keeps people from muting notifications, but it meant the loudest moment of the incident was the quietest. Missed blocks went from 689 to 5,989 during that silence.
So a halted node is now distinguished from one merely missing blocks: compare the rise in missed blocks against how far the chain moved. A ratio at or above 0.9 means we are signing nothing, and it lands in a single poll rather than waiting for a counter to look alarming. That alert repeats every poll for as long as the halt lasts, and a ladder at 25/50/75% of the jail threshold fires alongside it. Every alert now carries the projected jail time.