Quick Start
Get up and running with Refractr in 60 seconds. You'll need an API key: create an account, confirm your email address, and generate one from your dashboard.
import requests
BASE_URL = "https://api.refractr.io"
API_KEY = "re_your_api_key_here"
payload = {
"document_text": "Invoice #2847\nDate: January 15, 2026\nTotal: EUR 1,249.00",
"template": {
"invoice_number": "__str__",
"date": "__date__",
"total_amount": "__float__",
"currency": "__str__"
}
}
# Create the session once and reuse it for every call (keeps the connection open)
session = requests.Session()
session.headers["Authorization"] = f"Bearer {API_KEY}"
response = session.post(
f"{BASE_URL}/api/v1/extract/",
json=payload,
timeout=30
)
result = response.json()
print(result["extracted_data"])
# {"invoice_number": "2847", "date": "2026-01-15",
# "total_amount": 1249.0, "currency": "EUR"}
import base64
import requests
BASE_URL = "https://api.refractr.io"
API_KEY = "re_your_api_key_here"
# Send the file itself, base64-encoded
with open("invoice.pdf", "rb") as f:
pdf = base64.b64encode(f.read()).decode()
payload = {
"document": pdf,
"template": {
"invoice_number": "__str__",
"date": "__date__",
"total_amount": "__float__",
"currency": "__str__"
}
}
# Create the session once and reuse it for every call (keeps the connection open)
session = requests.Session()
session.headers["Authorization"] = f"Bearer {API_KEY}"
response = session.post(
f"{BASE_URL}/api/v1/extract/",
json=payload,
timeout=30
)
result = response.json()
print(result["extracted_data"])
# {"invoice_number": "2847", "date": "2026-01-15",
# "total_amount": 1249.0, "currency": "EUR"}
curl -X POST https://api.refractr.io/api/v1/extract/ \
-H "Authorization: Bearer re_your_api_key_here" \
-H "Content-Type: application/json" \
-d '{
"document_text": "Invoice #2847\nDate: January 15, 2026\nTotal: EUR 1,249.00",
"template": {
"invoice_number": "__str__",
"date": "__date__",
"total_amount": "__float__",
"currency": "__str__"
}
}'
# Put the file, base64-encoded, into the JSON body and send it
printf '{
"document": "%s",
"template": {
"invoice_number": "__str__",
"date": "__date__",
"total_amount": "__float__",
"currency": "__str__"
}
}' "$(base64 < invoice.pdf | tr -d '\n')" |
curl -X POST https://api.refractr.io/api/v1/extract/ \
-H "Authorization: Bearer re_your_api_key_here" \
-H "Content-Type: application/json" \
--data-binary @-
const response = await fetch(
"https://api.refractr.io/api/v1/extract/",
{
method: "POST",
headers: {
"Authorization": "Bearer re_your_api_key_here",
"Content-Type": "application/json"
},
body: JSON.stringify({
document_text: "Invoice #2847\nDate: January 15, 2026\nTotal: EUR 1,249.00",
template: {
invoice_number: "__str__",
date: "__date__",
total_amount: "__float__",
currency: "__str__"
}
})
}
);
const result = await response.json();
console.log(result.extracted_data);
import { readFile } from "node:fs/promises";
// Send the file itself, base64-encoded
const pdf = (await readFile("invoice.pdf")).toString("base64");
const response = await fetch(
"https://api.refractr.io/api/v1/extract/",
{
method: "POST",
headers: {
"Authorization": "Bearer re_your_api_key_here",
"Content-Type": "application/json"
},
body: JSON.stringify({
document: pdf,
template: {
invoice_number: "__str__",
date: "__date__",
total_amount: "__float__",
currency: "__str__"
}
})
}
);
const result = await response.json();
console.log(result.extracted_data);
The API returns your extracted data structured exactly as you defined it. Because each field declares its type, the values are guaranteed to be that type. Note the normalization: "January 15, 2026" comes back as "2026-01-15", and "EUR 1,249.00" becomes the bare number 1249.0 with the currency in its own field. See Templates & Types. To send a file such as a PDF or Word document instead of text, pick Documents above, or see Documents.
Authentication
All API requests must include your API key in the Authorization header.
Authorization: Bearer re_your_api_key_here
Keys work once you have confirmed your email address with the link we send at signup. Until then, requests return 403.
Templates & Types
A template is a JSON object shaped like the answer you want. Each leaf value declares the expected type using a typed placeholder. The type is enforced during generation, so the output is guaranteed to be that JSON type or null, not "usually". That guarantee is about shape, not accuracy: you always get well-formed, correctly typed JSON, but the values inside it come from a model and can be wrong.
| Placeholder | Output is always | Notes |
|---|---|---|
"__str__" | string or null | verbatim text span from the document |
"__int__" | integer or null | bare number: 3, never "3" or "3 items" |
"__float__" | number or null | bare: 4779.5, never "EUR 4,779.50" |
"__bool__" | true/false or null | never "yes"/"no" |
"__date__" | "YYYY-MM-DD" or null | ISO 8601, regardless of the format in the document |
Missing values
null means "not present in the document". Every typed field is nullable by design. A field is either a value of the declared type or null; never an empty string, a placeholder artifact, or malformed JSON.
Accuracy. Typed placeholders guarantee the shape of the response, not the correctness of the values. Extraction is model-based, so validate values against your own rules before treating them as authoritative, particularly for fields you are not certain the document contains.
Arrays
| Template form | Meaning |
|---|---|
["__str__"] | array of strings; an empty answer is [], never [null] or [""] |
[{"name": "__str__", "amount": "__float__"}] | array of objects with typed inner fields |
[] | untyped; accepts any array |
Untyped fields
A bare null leaf is still legal and means "any JSON type". Typed placeholders are recommended for every field where you know the type. They make results predictable and remove the need for defensive parsing on your side.
__something__ are reserved for placeholders. A template containing one that isn't in the table above (e.g. "__string__") is rejected with a 400 that names the offending field.
Canonical value formats
Rule of thumb: representation is normalized, content is verbatim. Refractr converts date and number formats to the canonical forms below, but never summarizes or paraphrases a text span.
| Semantic type | Format | Example |
|---|---|---|
| Date | "YYYY-MM-DD" | "2026-03-14" |
| Monetary amount | bare float, . decimal, no symbol | 4779.5 |
| Currency | separate field, ISO 4217 | "USD" |
| Quantity / count | bare integer | 3 |
| Boolean | JSON true/false | true |
| Missing scalar | null | |
| Empty list | [] | |
| Names / free text | verbatim from document | "Dr. Emily Watson" |
Documents
To extract from a file, such as a PDF, a Word document or a photo, send it as document, base64-encoded, in place of document_text. No file name or type is needed: Refractr detects the file type from the content. The template, the response and the price are the same as for text: 1 credit per extraction, or 2 with OCR.
import base64
with open("invoice.pdf", "rb") as f:
pdf = base64.b64encode(f.read()).decode()
payload = {
"document": pdf,
"template": {"invoice_number": "__str__", "total_amount": "__float__"}
}
Complete examples in Python, curl and JavaScript are in the Quick Start, under Documents.
Supported files
| Type | What is read |
|---|---|
| The text layer, which PDFs created by accounting, invoicing or office software have. With OCR, also scanned pages and text in images. | |
| Word and OpenDocument text: .docx, .odt | The text, including tables, headers, footers and footnotes. Pictures in the document are not read. |
| Images: JPEG, PNG, TIFF, WebP, BMP | With OCR only. Send iPhone photos (HEIC) as JPEG. |
| Plain text, such as .txt, .csv or an email | Send the text itself as document_text, not as a file. |
Other files are not supported, for example older Word files (.doc), RTF, spreadsheets and presentations. Save them as PDF first.
If you can select and copy the text in a PDF viewer, the PDF has a text layer. Without OCR, text inside images is not read, for example a scanned page or a photo pasted into the PDF, and a PDF without any text layer is rejected with 422 OCR_REQUIRED, not charged.
Word and OpenDocument files are read directly, so OCR does not apply to them: they cost 1 credit, also with "ocr": true. A document without any text, for example scans pasted in as pictures, is rejected with 422 NO_TEXT, not charged. Save it as PDF and send that with OCR instead.
Scans and photos (OCR)
For scans, photos and PDFs with scanned pages, add "ocr": true. Refractr then reads every page: from the text layer where it is complete, and with OCR where it is not. An extraction with OCR costs 2 credits, whatever the file contains, and usually adds well under a second.
with open("receipt.jpg", "rb") as f:
photo = base64.b64encode(f.read()).decode()
payload = {
"document": photo,
"template": {"merchant": "__str__", "total_amount": "__float__"},
"ocr": True
}
The response is the same as without OCR, with metadata.credits_charged set to 2. A file longer than the model can read in one go is extracted from its first part, and validation.warnings says so (DOCUMENT_TRUNCATED).
If OCR can't be done, the response has "status": "error" and no credits are charged. The error code is one of:
- OCR_NO_TEXT: no text was found, for example a blank page or a photo without text.
- OCR_FAILED: the file could not be read. Retrying may help.
- OCR_UNAVAILABLE: OCR is temporarily unavailable. Retry later.
Limits
| Limit | Value |
|---|---|
| File size | About 7 MB (requests are limited to 10 MB, and base64 encoding adds a third) |
| Pages | PDFs: up to 20, or 10 with OCR |
| Text | Up to 50,000 characters |
| Password-protected files | Not supported |
A file that is over a limit, of another type, password-protected or unreadable gets a 413 or 415 error, also not charged. See Error Codes.
Postman collection
Every endpoint, pre-filled with working examples, including a sample PDF, Word file and receipt image. Import it, paste your API key once, and send.
Setting it up
- In Postman, choose Import and select the downloaded file.
- Open the collection's Variables tab and paste your key into the current value of
api_key. Create a key from your dashboard; it works once your email address is confirmed. - Send Extract (sync). It ships with a sample invoice and a typed template.
The key is applied to every request as Authorization: Bearer {{api_key}} through collection-level auth, so you only set it once. base_url is a variable too, if you ever need to point it elsewhere.
The file requests, Extract a PDF, Extract a Word file and Extract with OCR, send small sample files kept in the collection variables sample_pdf, sample_docx and sample_receipt. To try your own file, replace a variable's current value with the file's base64.
job_id into a collection variable, so Poll job result works straight afterwards without copying anything by hand.
Scope & Accuracy
Refractr is built for one thing: simple-field extraction, fast and cheap. The sweet spot is a template of up to ~10 concrete fields (IDs, names, dates, amounts, booleans, short verbatim spans), extracted at around 500ms per document.
Out of scope
Fields that require derived reasoning or aggregation are not what this API is for: summaries, key-point lists, "all X mentioned in the document" sweeps, sentiment. The model will attempt them, but accuracy is materially lower; by design, this is not the product. If you need both, split your pipeline: use Refractr for the concrete fields and a separate step for derived ones.
invoice_date beats d1), and declare types for every field you can.
Reporting a bad extraction
Corrections from alpha users feed directly into model training. If an extraction comes back wrong, email [email protected] with the job_id from the response, the field in question, and either the correct value or "not in the document".
Both cases are useful, and the second is the one we can least easily determine on our own.
Extract Data
The core endpoint. Submit a document and get back structured data matching your template.
Request Body
{
"document_text": "BREAKING: Nvidia soars 12% to ~$187 after crushing Q4 earnings. Revenue hit $22.1B vs $20.4B expected.",
"template": {
"company": "__str__",
"stock_move_pct": "__float__",
"revenue_billions": "__float__",
"currency": "__str__"
},
"wait": true
}
Response (Sync Mode)
{
"job_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "success",
"extracted_data": {
"company": "Nvidia",
"stock_move_pct": 12.0,
"revenue_billions": 22.1,
"currency": "USD"
},
"confidence": {
"company": 0.9962,
"stock_move_pct": 0.9814,
"revenue_billions": 0.9731,
"currency": 0.8847
},
"validation": {
"warnings": []
},
"metadata": {
"model": "refractr-v9",
"inference_ms": 340,
"credits_charged": 1
}
}
validation.warnings lists informational notes about your template, for example that an empty string was normalized to null for a field. They're useful while iterating on a template and safe to ignore in production.
Confidence scores
Every successful extraction includes a confidence object next to extracted_data, at no extra cost. It has one entry per returned value: the estimated probability, from 0 to 1, that the value is correct.
| Key or score | Meaning |
|---|---|
line_items[0].amount | keys are value paths: dots for nesting, [i] for array items |
tags[0], tags[1] | a list of plain values has one key per item |
tags | an empty list has a single key for the list itself: the probability that there really is nothing to extract |
score on a null value | the probability that the field is really absent from the document |
null instead of a score | not scored: judgment fields (sentiment, category, type, summary, description, is_/has_/contains_ flags) and text values longer than eight words. Show these without a score rather than as 0. |
Refractr never filters or changes values based on confidence. You choose the threshold; a practical starting point is to review values below 0.9 first:
to_review = [path for path, score in result.get("confidence", {}).items()
if score is not None and score < 0.9]
If scores are unavailable for a request, the response has no confidence key and validation.warnings contains a CONFIDENCE_UNAVAILABLE entry with the reason. The extracted data is unaffected.
Response (Async Mode)
{
"job_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "queued",
"message": "Job submitted. Poll GET /api/v1/extract/{job_id}/ for results."
}
When wait: false, the API returns immediately with a job ID. Use the Poll Status endpoint to check when your result is ready.
503 rather than a hang. For spiky workloads or very large documents, prefer wait: false and poll.
requests.Session() instead of calling requests.post() directly; with httpx, create one httpx.Client(); in Node.js, fetch and most HTTP clients reuse connections automatically. This is free latency, with no code changes beyond creating the client once and reusing it.
Poll Status
Check the status of an async extraction job.
Response (Still Processing)
{
"job_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "processing",
"message": "Job is still processing.",
"queue_depth": 2
}
status is queued until a worker picks the job up, then processing. queue_depth is the number of jobs currently waiting ahead in the queue.
Response (Complete)
{
"job_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "success",
"extracted_data": { ... },
"confidence": { ... },
"validation": { "warnings": [] },
"metadata": { ... }
}
Pricing & Limits
Extractions are billed in credits, pay-as-you-go. Pricing is flat: it does not depend on document or template size.
| Item | Cost |
|---|---|
| Successful extraction | 1 credit |
| Successful extraction with OCR ("ocr": true) | 2 credits |
| Failed extraction (schema violation, timeout, worker error) | Free, no credit charged |
| Credit price | €1 = 200 credits (top up from €5) |
| Free daily allowance (alpha) | 100 free extractions per day |
The metadata.credits_charged field in each response tells you exactly what a call cost (0 for a failed extraction).
Rate limits
Requests are throttled per API key, 60 requests per minute by default. Exceeding the limit returns 429 with a Retry-After header; back off and retry after that interval.