# CSV reconcile

`csv-reconcile` · version 1.0.0 · CSV & tables · free, no key needed

Reconcile two CSV documents by key columns, reporting full population totals and bounded samples for added, removed, changed, invalid-key, and ambiguous rows.

**Use when you need to: csv reconcile · reconcile csv · csv reconciliation.**

## Supported

- csv reconcile
- reconcile csv
- csv reconciliation
- compare csv by keys with duplicates
- csv diff with duplicate keys
- keyed csv reconciliation report
- diagnose csv differences

## Not supported

- fuzzy matching
- numeric type coercion
- schema migration
- import or patch generation
- external writes
- filesystem access
- network access

## Behavior

- left and right must share the same exact set of header names; header order may differ.
- keys is a non-empty unique array of at most 16 column names existing in the headers.
- Empty string in any key cell makes that row invalid-key; whitespace is literal and non-empty.
- Invalid-key rows are strictly excluded from key matching and duplicate groups.
- Tuple identity uses collision-safe exact string tuples without trimming or case folding.
- Duplicate key (>1 valid-key rows on either side) quarantines all rows of that tuple on both sides under ambiguous_keys.
- Unambiguous valid tuples are classified into added (right-only), removed (left-only), changed (non-key diffs), or unchanged.
- changed.before and changed.after contain only changed columns, with changed_columns ordered according to left CSV header order.
- Full population conservation holds: left = invalid_left + ambiguous_left + removed + changed + unchanged; right = invalid_right + ambiguous_right + added + changed + unchanged.
- Samples are bounded per category by sample_limit (default 5, 0..20); limit 0 returns all counts and empty samples.
- Row numbers are 1-based data row indices.
- Added samples follow right row order; removed and changed follow left row order; invalid samples follow their own side row order; ambiguous keys follow first appearance in left, then unseen keys in right.
- Each category truncated flag equals samples.length < total; ambiguity side-row flags compare retained row indices with that side count.
- If JSON result exceeds 192 KiB, whole sample entries are deterministically shed by removing the last entry (retaining a first-seen prefix) in fixed category priority order (invalid_right, invalid_left, ambiguous_keys, added, removed, changed) while preserving exact counts.

## Input

- `left` (string, required): max length 102400
- `right` (string, required): max length 102400
- `keys` (array of string, required): min items 1; max items 16; each min length 1
- `sample_limit` (integer, optional): min 0; max 20; default 5

## Output

- `totals` (object, required)
- `added` (object, required)
- `removed` (object, required)
- `changed` (object, required)
- `invalid_left` (object, required)
- `invalid_right` (object, required)
- `ambiguous_keys` (object, required)

## Limits

- max input bytes: 245760
- max csv bytes: 102400
- max data rows: 5000
- max columns: 128
- max header name bytes: 256
- max keys: 16
- max sample limit: 20
- max output bytes: 196608

## Example

Request input:

```json
{
  "left": "id,name,city\n1,Ada,London\n2,Lin,Paris\n3,Zoe,Berlin\n4,Sam,Bristol\n5,Duplicate,CityA\n5,Duplicate,CityB\n,InvalidKey,NoId\n",
  "right": "city,name,id\nLondon,Ada,1\nRome,Lin,2\nOslo,Bob,6\nBristol,Sam,4\nCityA,Duplicate,5\nCityB,Duplicate,5\nNewCity,ValidKey,\n",
  "keys": [
    "id"
  ]
}
```

Response:

```json
{
  "result": {
    "totals": {
      "left_rows": 7,
      "right_rows": 7,
      "invalid_left": 1,
      "invalid_right": 1,
      "ambiguous_left": 2,
      "ambiguous_right": 2,
      "ambiguous_keys": 1,
      "added": 1,
      "removed": 1,
      "changed": 1,
      "unchanged": 2
    },
    "added": {
      "total": 1,
      "samples": [
        {
          "row": 3,
          "key": [
            "6"
          ],
          "values": {
            "city": "Oslo",
            "name": "Bob",
            "id": "6"
          }
        }
      ],
      "truncated": false
    },
    "removed": {
      "total": 1,
      "samples": [
        {
          "row": 3,
          "key": [
            "3"
          ],
          "values": {
            "id": "3",
            "name": "Zoe",
            "city": "Berlin"
          }
        }
      ],
      "truncated": false
    },
    "changed": {
      "total": 1,
      "samples": [
        {
          "key": [
            "2"
          ],
          "left_row": 2,
          "right_row": 2,
          "changed_columns": [
            "city"
          ],
          "before": {
            "city": "Paris"
          },
          "after": {
            "city": "Rome"
          }
        }
      ],
      "truncated": false
    },
    "invalid_left": {
      "total": 1,
      "samples": [
        {
          "row": 7,
          "key": [
            ""
          ],
          "values": {
            "id": "",
            "name": "InvalidKey",
            "city": "NoId"
          }
        }
      ],
      "truncated": false
    },
    "invalid_right": {
      "total": 1,
      "samples": [
        {
          "row": 7,
          "key": [
            ""
          ],
          "values": {
            "city": "NewCity",
            "name": "ValidKey",
            "id": ""
          }
        }
      ],
      "truncated": false
    },
    "ambiguous_keys": {
      "total": 1,
      "samples": [
        {
          "key": [
            "5"
          ],
          "tuple": [
            "5"
          ],
          "left_count": 2,
          "right_count": 2,
          "left_rows": [
            5,
            6
          ],
          "right_rows": [
            5,
            6
          ],
          "left_truncated": false,
          "right_truncated": false
        }
      ],
      "truncated": false
    }
  }
}
```

## How to call it

### MCP

Connect `https://computefirst.net/mcp` ([setup](/docs#connect)), then call `execute` with:

```json
{
  "id": "csv-reconcile",
  "version": "1.0.0",
  "input": {
    "left": "id,name,city\n1,Ada,London\n2,Lin,Paris\n3,Zoe,Berlin\n4,Sam,Bristol\n5,Duplicate,CityA\n5,Duplicate,CityB\n,InvalidKey,NoId\n",
    "right": "city,name,id\nLondon,Ada,1\nRome,Lin,2\nOslo,Bob,6\nBristol,Sam,4\nCityA,Duplicate,5\nCityB,Duplicate,5\nNewCity,ValidKey,\n",
    "keys": [
      "id"
    ]
  }
}
```

### HTTP (no key)

```sh
curl -X POST https://computefirst.net/v1/tools/csv-reconcile/versions/1.0.0/execute \
  -H "Content-Type: application/json" \
  -d '{"left":"id,name,city\n1,Ada,London\n2,Lin,Paris\n3,Zoe,Berlin\n4,Sam,Bristol\n5,Duplicate,CityA\n5,Duplicate,CityB\n,InvalidKey,NoId\n","right":"city,name,id\nLondon,Ada,1\nRome,Lin,2\nOslo,Bob,6\nBristol,Sam,4\nCityA,Duplicate,5\nCityB,Duplicate,5\nNewCity,ValidKey,\n","keys":["id"]}'
```

The machine-readable contract is at [/v1/tools/csv-reconcile/versions/1.0.0](/v1/tools/csv-reconcile/versions/1.0.0).

### CLI

```sh
node cli.mjs run csv-reconcile 1.0.0 --input input.json --base-url https://computefirst.net
```

Get the client at [/clients/cli/](/clients/cli/).

## Related tools

- [CSV row diff](/tools/csv-row-diff): Compare two CSV documents with identical headers by unique key tuples and report added, removed, and changed rows.
- [CSV dedupe](/tools/csv-dedupe): Remove duplicate CSV data rows while preserving the first occurrence.
- [CSV anti join](/tools/csv-anti-join): Keep left CSV rows whose exact key tuples do not appear in the right document.
- [CSV concat](/tools/csv-concat): Concatenate CSV documents that share identical headers, preserving document and row order.
- [CSV fill defaults](/tools/csv-fill-defaults): Replace empty cells in named CSV columns with exact default strings.
- [CSV filter equals](/tools/csv-filter-equals): Keep CSV data rows whose named column equals an exact string.
