Oct 5, 2026 · 4 min read
Ask an AI coding agent “what changed between these two JSON files?” and watch what it does. With small files you usually get a decent answer. With large ones — a minified API response, a 5,000-line fixture, a generated config — you typically get one of two failure modes: a wrong answer delivered with confidence, or a long, expensive hunt for the difference. The Skills and MCP integrations exist to fix both.
The first tool an agent reaches for is usually git diff --no-index (or plain diff). That’s a line-based text diff, and JSON is not line-based data.
The worst case is minified JSON — one line, tens of thousands of characters, extremely common in real API responses and snapshots:
$ git diff --no-index base.json contrast.json
-{"users":[{"id":1,"name":"Alice","email":"[email protected]"}, ... 40,000 more characters
+{"users":[{"id":1,"name":"Alice","email":"[email protected]"}, ... 40,000 more characters
One field changed somewhere in the middle, and the diff reports “the line changed”. The agent still has to locate the actual difference itself — inside a single 40 KB line. Some agents give up and guess, which is where confidently wrong answers come from.
Even with pretty-printed files, text diff produces misleading hunks: reordered keys, whitespace changes, or one element inserted into an array can cascade into dozens of “changed” lines that don’t correspond to any real data change (see what a single insertion does to By-Index-style alignment). A smarter agent will pretty-print both sides with jq first — better, but key order and array misalignment still generate noise. The agent faithfully summarizes that noise, and one change gets reported as a dozen.
The other default strategy is to read both files into the context window and look for the difference. On a large file this turns into a marathon: the file gets truncated, the agent re-reads it in chunks, runs a few greps, scrolls back and forth — and several minutes and tens of thousands of tokens later, it announces that meta.buildNumber went from 841 to 842.
Correct, but slow, expensive, and fragile. Every extra tool call is another chance to truncate, misalign, or miss a second difference hiding elsewhere in the file.
JSON diffing is a solved problem — just not by text tools. Parse both documents into trees, align objects by key and arrays by a chosen strategy, and report changes as exact paths. Deterministic, immune to formatting, and fast:
$ npx @compare-json/cli base.json contrast.json
valueChanged meta.buildNumber
One line of output, no matter how large the inputs. The agent’s job shrinks from finding the difference to explaining it — which is the part it’s actually good at. That’s the entire reason the skill and the MCP server exist: they put this engine one tool call away, so the comparison takes seconds instead of minutes, and the answer stops depending on how the files happened to be formatted.
Same engine, two delivery mechanisms:
The skill is a single SKILL.md that teaches your agent when and how to invoke @compare-json/cli. Install it once and any Agent Skills-compatible assistant — Claude Code, Codex CLI, OpenCode, Cursor — will run the CLI locally whenever you ask it to compare JSON:
npx skills add unitstack/compare-json
The MCP server exposes the same comparison as a structured compare_json tool. The CLI doubles as an MCP server (--mcp), so any MCP-capable client can call it with two file paths or strings and get machine-readable differences back:
{
"differences": [
{
"pathSegments": ["meta", "buildNumber"],
"pathString": "meta.buildNumber",
"pathBelongsTo": "both",
"diffType": "valueChanged"
}
]
}
Which one should you pick? If your assistant supports Agent Skills, the skill is the lighter touch — no server config, invoked on demand. If your client speaks MCP and you want a permanent, typed tool with structured output, use the MCP server. Both wrap the same engine, so they report identical differences — including the array comparison strategy and the case-insensitive options.
LLMs are good at reasoning about differences and bad at locating them in raw text. The skill and the MCP server exist to split that work properly: the engine finds every difference in seconds, deterministically, and your agent spends its time — and your tokens — on the part that needs judgment.