An invoice parser can return impeccable JSON and still charge the wrong order. The useful question is not “Did the model output JSON?” but “Which contract did the caller enforce, and what happens when the document cannot support a value?” Consider an invoice reading: Invoice INV-1042; PO PO-778; Total USD 19.00. Our application wants to propose a payable against an existing purchase order—not let a model approve a payment.
Choose the contract for the handoff
A plain prompt (“reply in JSON”) is a formatting request, not an output constraint. JSON mode constrains the response to valid JSON, but {"invoice_id":"INV-1042","total":"19.00"} is valid JSON while omitting the required po_id and currency. JSON mode is useful when a schema-constrained response is unavailable, provided the caller validates and rejects bad records. OpenAI also says JSON mode needs an explicit JSON instruction and that an incomplete response may not be parseable; do not treat every returned fragment as a finished record. OpenAI: JSON mode
For this extraction, a structured response is the natural handoff: the model returns data for the application to inspect. OpenAI's Structured Outputs can enforce adherence to a supplied schema on supported models and configurations, subject to its supported JSON Schema subset. If instead the model must request a call to application functionality, define a strict function tool with schema-constrained arguments. A tool call is a request to your code, not permission to run it. Validate and authorize before execution. Nonconforming or unsupported strict schemas can be rejected at request time; check the chosen model, API surface and schema against current documentation rather than assuming that arbitrary JSON Schema will work. OpenAI: Structured model outputs; OpenAI: function calling, strict mode
Make absence and interruption explicit
Use a small extraction contract, version invoice_extract_v1: required keys version, status, invoice_id, po_id, total, currency, evidence; no extra keys. status is ready or needs_review. The invoice ID, PO ID, total and currency are strings or null; evidence is a verbatim excerpt from the source, or null. For ready, all values and evidence must be present. If a field is absent or contradictory, return needs_review with null for the unresolved field and never invent a zero or a guessed currency. Use decimal strings for money, not binary floating-point numbers. This is an application contract; translate it into a schema supported by the selected provider before calling it. JSON Schema's required governs key presence, not truth or semantic correctness; a present null needs an allowed null type. JSON Schema: required properties
For example, a minimal schema for this handoff is below. It enforces keys, types and the status enum, but leaves the status-dependent rules and source matching to application code. This uses simple keywords commonly supported by strict structured output; confirm support on the actual API/model before deployment.
{"type":"object","properties":{
"version":{"type":"string","enum":["invoice_extract_v1"]},
"status":{"type":"string","enum":["ready","needs_review"]},
"invoice_id":{"type":["string","null"]},
"po_id":{"type":["string","null"]},
"total":{"type":["string","null"]},
"currency":{"type":["string","null"]},
"evidence":{"type":["string","null"]}
},"required":["version","status","invoice_id","po_id","total","currency","evidence"],"additionalProperties":false}
A provider-level refusal is not needs_review: handle the API's refusal signal without trying to parse it as an invoice. An incomplete response (including an output-limit interruption) is likewise not a complete extraction: stop, inspect the reason, and retry only when a bounded retry is appropriate. Transport errors and schema-request errors are separate states too. Only a completed, non-refused response goes through the extraction validator. The distinction is necessary even when the output format is strict. OpenAI: refusals and incomplete responses
See the gap between shape and truth
For the sample invoice, this is a complete-looking record with the right keys and types:
{"version":"invoice_extract_v1","status":"ready","invoice_id":"INV-1042","po_id":"PO-778","total":"119.00","currency":"USD","evidence":"Invoice INV-1042; PO PO-778; Total USD 19.00."}
Its total is wrong: 119.00 is nowhere in the excerpt. A schema can check that total is a string; it cannot establish that the model read the correct amount from this invoice. Even a source-matching amount could still conflict with the approved purchase order. The application must compare the proposed amount and order with its own trusted records, and must check that the requesting user is allowed to act on that order. Evidence is a review aid, not proof merely because the model quoted it.
A small downstream gate you can run
The following standard-library demonstration accepts an already completed, non-refused extraction. It is not an OpenAI API mock or a full JSON Schema validator: it checks this particular contract, a deliberately simple source pattern, an independently supplied order total and an authorization set. In production, validate the actual schema with a suitable validator and use a trusted invoice/document extraction and order system; do not rely on a substring check as fraud detection. Run the snippet with python invoice_gate.py after saving the code block.
import json
import re
from decimal import Decimal, InvalidOperation
SOURCE = "Invoice INV-1042; PO PO-778; Total USD 19.00."
KEYS = {"version", "status", "invoice_id", "po_id", "total", "currency", "evidence"}
GOOD = {"version": "invoice_extract_v1", "status": "ready",
"invoice_id": "INV-1042", "po_id": "PO-778", "total": "19.00",
"currency": "USD", "evidence": SOURCE}
def gate(raw, source, order_totals, authorized_pos):
def reject(reason):
return "REJECT: " + reason
try:
record = json.loads(raw)
except json.JSONDecodeError:
return reject("invalid JSON")
if not isinstance(record, dict) or set(record) != KEYS:
return reject("contract keys")
if record["version"] != "invoice_extract_v1":
return reject("unknown version")
if record["status"] not in ("ready", "needs_review"):
return reject("status")
fields = ("invoice_id", "po_id", "total", "currency", "evidence")
if any(value is not None and not isinstance(value, str)
for value in (record[key] for key in fields)):
return reject("field type")
if record["status"] == "needs_review":
return "REVIEW: unresolved extraction; no action"
if any(not record[key] for key in fields):
return reject("ready requires all values")
if record["evidence"] not in source:
return reject("evidence not in source")
# Demonstration grammar for this one sample, not general invoice OCR.
match = re.fullmatch(
r"Invoice (INV-\d+); PO (PO-\d+); Total (USD) (\d+\.\d{2})\.",
record["evidence"])
if not match or tuple(record[k] for k in
("invoice_id", "po_id", "currency", "total")) != match.groups():
return reject("values do not match source")
try:
amount = Decimal(record["total"])
except InvalidOperation:
return reject("amount")
po = record["po_id"]
if po not in order_totals or (record["currency"], amount) != order_totals[po]:
return reject("order amount/currency mismatch")
if po not in authorized_pos:
return reject("user not authorized for order")
return "PASS: proposal may enter approval; do not pay automatically"
def run(label, record, authorized={"PO-778"}):
print(label, gate(json.dumps(record), SOURCE,
{"PO-778": ("USD", Decimal("19.00"))}, authorized))
if __name__ == "__main__":
run("correct", GOOD)
run("missing key", {k: v for k, v in GOOD.items() if k != "currency"})
run("wrong total", {**GOOD, "total": "119.00"})
run("unauthorized", GOOD, set())
run("unknown", {**GOOD, "status": "needs_review", "total": None})
The gate rejects the syntactically valid missing-key record, rejects the schema-shaped but ungrounded amount, stops an unauthorized request, and routes the unknown amount to review. The PASS outcome only permits an approval workflow; it never triggers payment.
Route failures, then version the tests
| Observed state | Caller action |
|---|---|
| Transient transport failure | Bounded retry with idempotency protection; no duplicate action. |
| Invalid JSON or wrong keys on a completed response | Reject; at most one corrected extraction attempt, then review. Investigate configuration if using strict output. |
| Refusal or incomplete response | Follow the API's refusal/incomplete branch; never deserialize as a finished invoice. Retry an interruption only if its cause is remediable. |
| Missing or conflicting source field | Request a better document or human review, not repeated guessing. |
| Source mismatch, order mismatch or unauthorized order | Block action and escalate; a new model answer cannot grant authority. |
| Unsupported strict schema/request error | Fix the schema or choose a supported model/configuration; do not retry the identical request. |
Pin the extraction contract version alongside model identifier, prompt version and validator version. Reject unknown versions instead of silently interpreting changed fields. Regression cases should include the five snippet outcomes plus absent currency, a contradictory total, another currency, duplicate invoice identifiers, and a truncated/refused API response (tested at the API boundary, not by this local gate). Count failures by category and review actual invoice evidence before widening automation. For a related question about action authority in shopping flows, see Meydo Journal's article on AI-agent shopping payments.
