Verifiable analysis of engine test data
Every number in the report has to come from a computation.
Groundline is an AI agent for rocket-engine test data. The language model never does arithmetic: deterministic tools compute every value and log it in an evidence ledger. A verifier then checks every number the model writes: it must appear in the evidence the claim cites, with a unit and a meaning that fit the field it came from.
Below is that verifier running in your browser on real evidence from a public solid-motor static fire. Try to get a wrong number past it.
uvx groundline mcpLive · runs in your browser
Try the verifier
Evidence: SNU Rocket Team HANARO, KNSB static fire, analysed by Groundline's tools (E1–E5 on the right). Pick an example or edit the claim; every number is re-checked as you type.
Claim
1.0 found1.0 found, wrong meaning1.0 not in evidenceEvidence ledger
computed by the tools, not the modelHow it works
One analysis, in the order it runs.
- Load the runCSV or TDMS plus channel metadata, redlines and an optional simulation prediction.
- Tools computePhase segmentation, sensor health, redlines, valve response, oscillation FFT, impulse. No model involved.
- Ledger recordsEach call becomes E1, E2, … with parameters, data SHA-256, tool source hash and result.
- Model writesThe agent chooses tools and writes findings that may only cite ledger entries.
- Verifier checksEvery number is traced to a field; failures go back to the model once, then get flagged in the report.
Results
Raw results are in the repository under docs/results/.
Models on the same 14 synthetic hot-fires
| Model | Reports | Recall (strict / loose) | Precision | First-draft numbers with no source |
|---|---|---|---|---|
| Rule agent (no LLM) | 14/14 | 100% / 100% | 100% | – |
| gpt-5.6-sol | 14/14 | 100% / 100% | 95% | 0 / 376 (0%) |
| Qwen2.5 7B, local | 11/14 | 5% / 50% | 5% | 23 / 98 (23%) + 4 wrong meaning |
| Qwen2.5 3B, local | 13/14 | 0% / 0% | – | 1 / 1 |
Every number the verifier stopped in the 7B drafts was checked by hand against recomputed evidence: invented values and values cited from the wrong evidence, no false alarms. Small samples and large run-to-run variance: a rough guide, not a ranking.
The verifier's own test: errors planted in correct findings
| Planted error | Count | Grounding only | Grounding + meaning |
|---|---|---|---|
| Invented value (−30% to +50%) | 1,364 | 87% | 95% |
| Real value in the wrong place | 1,436 | 0% | 45% |
| Wrong unit (s ↔ Hz, % → s) | 1,168 | 0% | 92% |
| Unchanged correct findings (false alarms) | 268 | 0% | 0% |
A value swapped for another field of the same kind (one time for another) is caught only 13% of the time, when a role word such as "peak" or "lasting" gives it away. That is the limit of checking text against field names.
Real data: HANARO static fire against the team's own processing
| Groundline | HANARO processed | |
|---|---|---|
| Peak thrust | 2222.2 N | 2221.7 N |
| Total impulse | 6391 N·s (2% threshold, 4.12 s) | 6411 N·s (4.35 s window) |
Thrust calibration was fitted to HANARO's curve, so matching peaks are expected; clock alignment, action time and impulse are the independent checks. Open the full report.
Use it
Runs locally; test data never leaves the machine. The model only sees tool summaries.
Claude Code
claude mcp add --scope user groundline -- uvx groundline mcpClaude Desktop, Cursor, other MCP clients
{"mcpServers": {"groundline": {
"command": "uvx", "args": ["groundline", "mcp"]}}}Command line
pip install groundline
groundline demo
groundline analyze run.csv --agent openaiBenchmarks
groundline leaderboard examples/leaderboard.json
groundline verifier-bench --n 50