Groundlinev0.1.1

Verifiable analysis of engine test data

Every number in the report has to come from a computation.

Groundline is an AI agent for rocket-engine test data. The language model never does arithmetic: deterministic tools compute every value and log it in an evidence ledger. A verifier then checks every number the model writes: it must appear in the evidence the claim cites, with a unit and a meaning that fit the field it came from.

Below is that verifier running in your browser on real evidence from a public solid-motor static fire. Try to get a wrong number past it.

uvx groundline mcp

Live · runs in your browser

Try the verifier

Evidence: SNU Rocket Team HANARO, KNSB static fire, analysed by Groundline's tools (E1–E5 on the right). Pick an example or edit the claim; every number is re-checked as you type.

Claim

1.0 found1.0 found, wrong meaning1.0 not in evidence
Cites

    Evidence ledger

    computed by the tools, not the model

    How it works

    One analysis, in the order it runs.

    1. Load the runCSV or TDMS plus channel metadata, redlines and an optional simulation prediction.
    2. Tools computePhase segmentation, sensor health, redlines, valve response, oscillation FFT, impulse. No model involved.
    3. Ledger recordsEach call becomes E1, E2, … with parameters, data SHA-256, tool source hash and result.
    4. Model writesThe agent chooses tools and writes findings that may only cite ledger entries.
    5. Verifier checksEvery number is traced to a field; failures go back to the model once, then get flagged in the report.

    Results

    Raw results are in the repository under docs/results/.

    Models on the same 14 synthetic hot-fires

    ModelReportsRecall (strict / loose)PrecisionFirst-draft numbers with no source
    Rule agent (no LLM)14/14100% / 100%100%–
    gpt-5.6-sol14/14100% / 100%95%0 / 376 (0%)
    Qwen2.5 7B, local11/145% / 50%5%23 / 98 (23%) + 4 wrong meaning
    Qwen2.5 3B, local13/140% / 0%–1 / 1

    Every number the verifier stopped in the 7B drafts was checked by hand against recomputed evidence: invented values and values cited from the wrong evidence, no false alarms. Small samples and large run-to-run variance: a rough guide, not a ranking.

    The verifier's own test: errors planted in correct findings

    Planted errorCountGrounding onlyGrounding + meaning
    Invented value (−30% to +50%)1,36487%95%
    Real value in the wrong place1,4360%45%
    Wrong unit (s ↔ Hz, % → s)1,1680%92%
    Unchanged correct findings (false alarms)2680%0%

    A value swapped for another field of the same kind (one time for another) is caught only 13% of the time, when a role word such as "peak" or "lasting" gives it away. That is the limit of checking text against field names.

    Real data: HANARO static fire against the team's own processing

    GroundlineHANARO processed
    Peak thrust2222.2 N2221.7 N
    Total impulse6391 N·s (2% threshold, 4.12 s)6411 N·s (4.35 s window)

    Thrust calibration was fitted to HANARO's curve, so matching peaks are expected; clock alignment, action time and impulse are the independent checks. Open the full report.

    Use it

    Runs locally; test data never leaves the machine. The model only sees tool summaries.

    Claude Code

    claude mcp add --scope user groundline -- uvx groundline mcp

    Claude Desktop, Cursor, other MCP clients

    {"mcpServers": {"groundline": {
      "command": "uvx", "args": ["groundline", "mcp"]}}}

    Command line

    pip install groundline
    groundline demo
    groundline analyze run.csv --agent openai

    Benchmarks

    groundline leaderboard examples/leaderboard.json
    groundline verifier-bench --n 50