CSV dedupe
csv-dedupe · version 1.0.0 · CSV & tables · free, no key needed
Remove duplicate CSV data rows while preserving the first occurrence.
Use when you need to: csv dedupe · csv deduplicate · remove duplicate csv rows.
Supported
- csv dedupe
- csv deduplicate
- remove duplicate csv rows
- dedupe csv by named columns
Not supported
- normalize values
- fuzzy matching
- merge rows
- read files
- fetch urls
Behavior
- The first CSV record is the header row.
- Duplicate and empty header names are rejected.
- Without key_columns, equality uses every field in row order.
- With key_columns, equality uses the exact string values of those named columns.
- No trimming, case folding, type coercion, or other normalization is performed.
- The first data row for each key is retained; output order is unchanged.
Input
csv(string, required): max length 256000key_columns(array of string, optional): min items 1; max items 100; each min length 1
Output
csv(string, required)rows_in(integer, required)rows_out(integer, required)duplicates_removed(integer, required)
Limits
- max input bytes: 256000
- max data rows: 5000
- max columns: 256
- max header name bytes: 256
- max key columns: 100
Example
Request input:
{
"csv": "id,name\n1,Ada\n1,Ada\n2,Lin\n"
}
Response:
{
"result": {
"csv": "id,name\n1,Ada\n2,Lin\n",
"rows_in": 3,
"rows_out": 2,
"duplicates_removed": 1
}
}
How to call it
MCP
Connect https://computefirst.net/mcp (setup), then call execute with:
{
"id": "csv-dedupe",
"version": "1.0.0",
"input": {
"csv": "id,name\n1,Ada\n1,Ada\n2,Lin\n"
}
}
HTTP (no key)
curl -X POST https://computefirst.net/v1/tools/csv-dedupe/versions/1.0.0/execute \
-H "Content-Type: application/json" \
-d '{"csv":"id,name\n1,Ada\n1,Ada\n2,Lin\n"}'
The machine-readable contract is at /v1/tools/csv-dedupe/versions/1.0.0.
CLI
node cli.mjs run csv-dedupe 1.0.0 --input input.json --base-url https://computefirst.net
Get the client at /clients/cli/.
Related tools
- CSV group count: Count CSV rows by exact named-column tuples in first-seen order.
- CSV sort: Stably sort CSV rows by exact named string columns.
- CSV split column: Split one CSV column on an exact separator into uniquely named columns.
- CSV drop columns: Drop named CSV columns and keep the remaining columns in original header order.
- CSV reconcile: Reconcile two CSV documents by key columns, reporting full population totals and bounded samples for added, removed, changed, invalid-key, and ambiguous rows.
- CSV row diff: Compare two CSV documents with identical headers by unique key tuples and report added, removed, and changed rows.