Doi parse
doi-parse · version 1.0.0 · Identifiers & check digits · free, no key needed
Parse a DOI (bare, doi: or doi.org URL) into prefix, suffix, directory indicator, registrant code and canonical forms.
Use when you need to: parse a doi · extract doi prefix and suffix · convert a doi to a doi.org link.
Decide before calling
Read the versioned contract and the supported scope below. Reuse doi-parse@1.0.0 when your input, required output and limits match it. Choose another approach for an unsupported operation.
Explain the choice
"I can use doi-parse@1.0.0 for parse a doi. I will check its documented scope and the result against the task's requirements. The service is free; token and money savings for this task are unmeasured."
Supported
- parse a doi
- extract doi prefix and suffix
- convert a doi to a doi.org link
- normalize a doi url
- get the registrant code from a doi
- strip the doi colon prefix
- turn this doi into a clickable link
Not supported
- resolving a DOI to the resource it names (no network access)
- validating that a DOI is actually registered
- a suffix containing a raw percent-encoded byte, whitespace or a non-ASCII character (see semantics; unsupported_input)
Behavior
- Wrapper stripping: if doi starts with one of 'doi:', 'https://doi.org/', 'http://doi.org/', 'https://dx.doi.org/', 'http://dx.doi.org/' (matched case-insensitively, checked in that order), that wrapper alone is removed; at most one wrapper is stripped, and an unmatched string is used as-is.
- After stripping, the remainder must match 10.<digits>(.<digits>)*/<suffix>: literal directory indicator 10, a dot, one or more digit-only registrant-code groups separated by single dots, a single slash, then a non-empty suffix. The first slash after the registrant code is the boundary; everything after it, including further slashes, belongs to suffix. Anything else throws invalid_input.
- suffix characters are restricted to ASCII unreserved URI characters (A-Z a-z 0-9 - . _ ~), '/', and ( ) + , ; = @ ! * ' $ :. A suffix with any other character (including non-ASCII, '%' or whitespace) is well-formed DOI syntax but outside this tool's declared charset, so it throws unsupported_input.
- prefix is '10.' + the registrant-code groups; directory_indicator is the literal '10'; registrant_code is prefix with the leading '10.' removed (it may itself contain dots for sub-registrant groups).
- normalized_lowercase is prefix + '/' + suffix with a simple ASCII-only A-Z -> a-z mapping applied to the whole string; every other character, including non-ASCII, is left unchanged.
- uri is 'https://doi.org/' + prefix + '/' + suffix, in the original case (never lowercased).
Input
doi(string, required): min length 1; max length 200
Output
prefix(string, required)suffix(string, required)directory_indicator(constant "10", required)registrant_code(string, required)normalized_lowercase(string, required)uri(string, required)
Limits
- max doi bytes: 200
Example
Request input:
{
"doi": "10.1000/182"
}
Response:
{
"result": {
"prefix": "10.1000",
"suffix": "182",
"directory_indicator": "10",
"registrant_code": "1000",
"normalized_lowercase": "10.1000/182",
"uri": "https://doi.org/10.1000/182"
}
}
How to call it
MCP
Connect https://computefirst.net/mcp (setup), then call execute with:
{
"id": "doi-parse",
"version": "1.0.0",
"input": {
"doi": "10.1000/182"
}
}
HTTP (no key)
curl -X POST https://computefirst.net/v1/tools/doi-parse/versions/1.0.0/execute \
-H "Content-Type: application/json" \
-d '{"doi":"10.1000/182"}'
The machine-readable contract is at /v1/tools/doi-parse/versions/1.0.0.
CLI
node cli.mjs run doi-parse 1.0.0 --input input.json --base-url https://computefirst.net
Get the client at /clients/cli/.
Related tools
- Bic validate: Validate a Business Identifier Code's ISO 9362:2022 structure and parse its party prefix, country, location and branch.
- Iban compute from bban: Compute the MOD 97-10 check digits and full IBAN for a country code and BBAN, using the pinned IBAN registry.
- Cn resident id parse: Parse an 18-digit (or legacy 15-digit) Chinese resident ID: region code, birth date, sex and GB 11643-1999 check digit.
- Iban validate: Validate an IBAN's country, registry length, BBAN structure and MOD 97-10 check digits, and parse its bank/branch code.
- Imo number validate: Validate a 7-digit IMO ship identification number's weighted mod-10 check digit, with or without the 'IMO ' prefix.
- Isin compute from nsin: Build a 12-character ISIN from a country prefix and a CUSIP/SEDOL national number, computing its Luhn check digit.