Map an External Data Source
@monospace/cli 0.2 and @monospace/extension-kit 0.3. Overview
Use this page to design a data connector extension for a new external data source (an API, a service, or a database). The connector exposes it as collections in a Monospace workspace, and each item in a collection is one object from the external data source.
Work through the sections in order: probe the external data source, write a capability table, choose a read path for its query grammar, then handle paging, keys, writes, values and errors. You end with a declaration that offers only what the external data source can answer, and a read path that serves all of it.
Two connectors illustrate the patterns:
| Connector | Query grammar | Used in |
|---|---|---|
Art Institute of Chicago (artworks, artists) | Rich: an Elasticsearch query body on one search endpoint | This page. The Data Connector Quickstart builds a trimmed version that declares only equals, in and and, a sort on id only, and no count |
Stripe test mode (customers, products, prices) | Weak: list endpoints with a few exact parameters | Build a Stripe Data Connector |
Probe the External Data Source
An API's reference is wrong in both directions. Stripe's list endpoints accept an offset parameter its reference never mentions, and the Art Institute's search stops at 1,000 results, not the 10,000 its documentation states. Call every capability you might rely on, and record what comes back.
| Area | What to probe | Example finding |
|---|---|---|
| Metadata | How collections, fields, keys and references are described. Without a metadata API, the schema is static | The Stripe connector declares its schema statically, and maps each declared field to its place in a Stripe object |
| Filters | Each parameter with a value that must match nothing. Items coming back mean the parameter is ignored | Stripe rejects unknown parameters with an error instead of ignoring them |
| Exactness | Whether a filter matches exactly or returns a superset, and at which precision ranges compare | Stripe's email filter is exact and case-sensitive. The Art Institute's title keywords are lowercased, so a term query returns a superset of an exact match |
| Unknown values | An unknown enum value or a reference to a missing object: empty result or error | products?active=maybe and prices?product=prod_does_not_exist both fail on Stripe |
| Nulls | How the external data source tests for a missing value, and whether its negations include nulls | Elasticsearch must_not includes documents where the field is missing, which SQL NOT excludes |
| Sort | Which fields sort in which directions, and where nulls land | Stripe lists only newest created first. The Art Institute sorts id and dateStart both ways, with nulls placed by missing. Its title sort uses the lowercased keyword |
| Paging | Offset, cursor or both; the largest page; any maximum depth | Stripe caps a page at 100 without an error. The Art Institute refuses size above 100 and any request past result 1,000 |
| Batches | Whether a lookup by several ids keeps order and reports missing ids | The Art Institute's ids parameter reorders results and drops missing ids without notice |
| Counts | Whether the external data source returns an exact total | The Art Institute returns pagination.total. Stripe's list endpoints return no total |
| Writes | What can be created, updated and deleted, what each call returns, and what can never be deleted | Stripe answers a delete with a tombstone, can't delete prices, and refuses to delete a product that has prices |
| Limits | Rate limits, what a test run costs against them, and whether you can pin an API version | Stripe pins versions with Stripe-Version. The Art Institute's documentation states 60 requests per minute per IP without a key |
| Errors | What a bad credential, a missing object and invalid input return | Stripe answers a wrong key with 401 and a missing object with resource_missing |
Without credentials, read the vendor's OpenAPI or GraphQL schema and probe the unauthenticated endpoints, including the 401 path.
Write a Capability Table
List every filter operator, sort, paging argument and count you might offer callers, and name the request to the external data source that answers each one. If nothing answers it, don't offer it.
This excerpt is the Art Institute connector's table for artworks:
| Offered | Request to the Art Institute API | Evidence |
|---|---|---|
id: equals, in | term / terms on id | terms on 3 ids, one missing, returns exactly the 2 that exist |
artistId: equals, in | term / terms on artist_id | Count matches the Art Institute API's own count for the same artist |
dateStart: equals, in, ranges | term / terms / range on date_start | gt 1884 and gte 1885 return the same total on this integer field |
isPublicDomain: equals | term on is_public_domain | true plus false equals the unfiltered total |
sort id, dateStart, both directions | sort with missing for nulls | Nulls first or last as asked. Without nulls, last in both directions, where the search API puts them by default |
meta.totalCount | the same query with size: 0 | pagination.total |
Fields the Art Institute API can't match exactly stay out of the table. The Art Institute connector declares no operator on title, because its keyword matching is case-insensitive and equals would return a superset.
Answering a filter by fetching everything and filtering in memory turns one Studio click into a scan of the whole external data source. Leave it out of the declaration instead. The declaration rules, and which operators each field type allows, are in Operations and Operators.
Match the Read Path to the Grammar
The external data source's query grammar decides the shape of your read path.
| Grammar | The external data source accepts | Your connector |
|---|---|---|
| Rich | A filter language of its own (e.g., Elasticsearch, GraphQL filters, a query DSL) | Translates the filter tree structurally into one query per read |
| Weak | A few list parameters with fixed meanings | Plans the filter into a union of exact requests to the external data source, or refuses it |
Translate a Rich Grammar
The Art Institute connector maps each filter node to an Elasticsearch clause. Every read sends the translated query, in pages of up to 100 items, after a count request when the read reaches past result 1,000. The snippets on this page are excerpts from the full example sources. A file's first excerpt starts with its imports, and later excerpts of the same file continue it; // … marks the lines an excerpt leaves out. The Data Connector Quickstart builds a trimmed Art Institute connector: it declares no ranges, or, or counts, sorts only on id, and leaves out the declaration and value checks. Build a Stripe Data Connector walks through the Stripe connector:
import type {
DeclaredOperator,
FilterContract,
FilterSerialized,
LogicalOperator,
MonospaceObject,
PrimitiveFieldType,
QueryResultFor,
QueryResultSerialized,
ReadArguments,
ReadManyOperationSerialized,
ReadOneOperationSerialized,
SortItemSerialized,
Value,
} from '@monospace/extension-kit/data-connector';
import { ConnectionFailed, QueryRejected, RateLimited } from '@monospace/extension-kit/data-connector';
import { schema } from './schema';
// …
/**
* Translates a filter into an Elasticsearch query. A field or operator `contract` doesn't declare,
* anywhere in the filter, is refused with a plain error, never dropped: it can be a permission rule.
* A shape the API can't answer, such as an `or` of permission rules on a collection that doesn't
* declare `or`, is a `QueryRejected`.
*/
function toQuery(filter: FilterSerialized, contract: FilterContract | undefined): SearchBody {
requireDeclared(filter, contract);
return translate(filter, contract);
}
function translate(filter: FilterSerialized, contract: FilterContract | undefined): SearchBody {
if (filter.and !== undefined) {
requireLogical(contract, 'and');
return { bool: { filter: filter.and.map((child) => translate(child, contract)) } };
}
if (filter.or !== undefined) {
requireLogical(contract, 'or');
return { bool: { should: filter.or.map((child) => translate(child, contract)), minimum_should_match: 1 } };
}
if (filter.not !== undefined) {
// Nothing declares `not`, so Engine offers no negation, and no `_null`, which needs it.
throw new QueryRejected('The Art Institute connector can\'t negate a filter.');
}
if (filter.in !== undefined) {
if (filter.in.field === undefined || !Array.isArray(filter.values)) throw new QueryRejected('`in` needs a field and a list of values.');
// The search API refuses `null` in `terms`, and `terms` never matches a missing value. So null
// members are left out, and `in` never matches an item without a value.
return { terms: { [apiFieldName(filter.in.field)]: filter.values.filter((value) => value !== null) } };
}
const { field } = filter.left;
const { value } = filter.right;
if (field === undefined || value === undefined) {
throw new QueryRejected('The Art Institute connector needs a field on the left of a comparison and a value on the right.');
}
const name = apiFieldName(field);
if (value === null) {
// `field = null` means "is null": the search API's `exists` is false for a field without a value.
if (filter.cmp === 'equals') return { bool: { must_not: { exists: { field: name } } } };
// Any other comparison with null is unknown in SQL, so it matches nothing.
return { match_none: {} };
}
switch (filter.cmp) {
case 'equals': return { term: { [name]: value } };
case 'lessThan': return { range: { [name]: { lt: value } } };
case 'lessThanOrEquals': return { range: { [name]: { lte: value } } };
case 'greaterThan': return { range: { [name]: { gt: value } } };
case 'greaterThanOrEquals': return { range: { [name]: { gte: value } } };
default: throw new QueryRejected(`The Art Institute connector can't translate \`${filter.cmp}\`.`);
}
}
// …
/** Checks every comparison in the filter, however deeply nested, before any of it is translated. */
function requireDeclared(filter: FilterSerialized, contract: FilterContract | undefined): void {
if (filter.cmp !== undefined) requireOperator(contract, filter.left.field, filter.cmp);
if (filter.in !== undefined) requireOperator(contract, filter.in.field, 'in');
for (const child of filter.and ?? filter.or ?? (filter.not === undefined ? [] : [filter.not])) requireDeclared(child, contract);
}
/** Refuses `operator` on `field` unless `contract` declares it. A value in the field's place is a shape `translate` refuses. */
function requireOperator(contract: FilterContract | undefined, field: string | undefined, operator: DeclaredOperator): void {
if (field === undefined) return;
const operators = contract?.fields[field]?.operators;
if (operators === undefined || !Object.hasOwn(operators, operator)) {
throw new Error(`The Art Institute connector can't filter \`${field}\` with \`${operator}\`.`);
}
}
- Only what the collection declares is translated.
requireDeclaredchecks every comparison in the filter, however deeply nested, against the connector's own declaration. It checks the field on the left of each comparison and the field of eachin. An undeclared field or operator there is a plainError, thrown before anything is translated or sent (e.g., "The Art Institute connector can't filterartistIdwithlessThan."). A field on the right of a comparison is a shapetranslaterefuses withQueryRejected. - Engine doesn't check comparisons for you. Engine checks a caller's request against the declaration it stored for the data source, but sends its own filters unchecked. An undeclared comparison reaches
requireDeclared, for example, when that stored declaration is older than the code, after an extension update that narrowed the declaration without a schema refresh. On a collection with writes, a permission rule's condition can also bring one. See Handle Requests You Didn't Declare. - Refuse, never drop. A comparison you skip can be a permission rule, and dropping it widens what the caller reads. No class fits an undeclared comparison, and Engine reports the plain
Erroras a failed query.requireLogicalthrowsQueryRejectedinstead, because Engine composesoracross permission rules whatever the collection declares. andandormap directly.bool.filterandbool.shouldwithminimum_should_match: 1.equalswithnullneeds a null test. It means "is null", never a literal. The connector sendsmust_not exists.- Other comparisons with
nullmatch nothing. In SQL they're unknown, so the connector sendsmatch_none. inleavesnullmembers out. The search API refusesnullinterms, andtermsnever matches a missing value. Soinnever matches an item without a value.notis rejected. The connector doesn't declarenot, because Elasticsearch'smust_notselects missing values that SQLNOTexcludes. So Engine offers callers no negation, and no_null, which needsnot. Anotthat still arrives is aQueryRejectedthat says so: "The Art Institute connector can't negate a filter."
When an external data source has no general negation, push not down to the leaves (and becomes or, equals becomes not-equals, < becomes >=). Refuse a leaf with no complement.
How a negated leaf treats a null field depends on the semantics you implement: your connector defines its own null handling. Under SQL three-valued logic, which the example connectors follow, a negated comparison never matches a null field. There, add an explicit "is not null" next to each negated comparison on a nullable field. Document the rules you choose, because callers can't read them from the declaration.
Plan a Weak Grammar
The Stripe connector plans each filter into Stripe requests before it sends any:
andintersects the allowed values and ranges of its children.inandorfan out into one list request per value or branch. A filter that plans more than 250 requests is refused. The bound counts planned requests, not HTTP requests, since one planned list can take several pages. It's checked while the combinations expand, before duplicates merge, so a filter can be refused even when fewer distinct requests would remain.- An id constraint becomes a retrieve by id, one request per id. Top-level ids bypass the 250 bound; the read budget in Bound Unbounded Reads limits them instead.
- Anything else is refused. A field or operator the collection doesn't declare is a plain
Error. A permission rule's condition or an extension update without a schema refresh can send one. A declared shape Stripe can't express, such asnotorfield = null, is aQueryRejected: "Stripe can't filtercustomersbyemailbeing null without reading every item."
The plan then runs in one of two shapes:
| Plan | When | Paging |
|---|---|---|
| Page | The filter is exactly one list request | offset and limit pass through to Stripe. Engine's per-parent relation reads (in: [key]) land here |
| Collect | Several list requests or retrieves by id | Each list request reads offset + limit items, plus any created in the same second as the last one, and each retrieve fetches its ids. The connector drops duplicate ids, merges the rest in the declared order, re-checks the filter, and slices |
The two plans can order items created in the same second differently. The page plan keeps Stripe's own order for such ties, which Stripe doesn't document. The collect plan breaks them by id.
Build a Stripe Data Connector walks through the key parts of this code, and its downloadable project has all of it.
Apply These Rules to Both Grammars
- Read both operands. Engine can send
value < fieldas well asfield > value. Reverse the ordered comparisons, and refuse string operators on a reversed pair. - Refuse field-to-field comparisons unless the external data source expresses them. Never send a field name as a literal.
- Convert ranges to the external data source's precision, exactly. Stripe stores
createdin whole seconds, so> tbecomescreated[gte]=floor(t)+1. Parse the timestamp strictly and keep every fractional digit.>=a timestamp one microsecond past a whole second (…:37.000001Z) must move the bound to the next second, and a parser that stops at the millisecond keeps it on the same one. - Decide out-of-domain values locally. Stripe rejects
active=maybewith an error. The Stripe connector never sends it, because no item can hold that value. - Never turn "no match" into an error. Stripe answers
prices?product=with a missing product as an error. The Stripe connector answers it with no items. - Refuse sort options you can't honor, such as explicit null placement on an external data source that has none, rather than ignoring them.
- Keep a fixed null placement when
nullsis absent. Engine sendsnullsonly when the caller sets it. Without it, your connector picks the placement. Pick one, such as where the external data source puts them by default, and keep it stable across pages. The Art Institute connector sendsmissing: '_last'in both directions, where the search API puts items without a value by default.
Re-Check Every Item
Run the full filter over every item the external data source returns, unless the request is exact by construction. This residual check makes a superset parameter correct when you page over the items that pass it: skip the caller's offset among those items, then keep reading until limit of them pass. Never forward the caller's offset to a request that returns a superset, because the items it skips aren't all matches. The check can't recover an item the request excluded, so pushdown must never drop a match.
The Stripe connector evaluates with SQL three-valued logic. That's the examples' choice, not a kit rule:
| Case | Result |
|---|---|
A comparison with null | Unknown |
equals with a null literal | True when the field is null. This is how Engine sends "is null" |
not of unknown | Unknown |
in with no matching member and a null member | Unknown, so not around it stays unknown |
and with any false child | False; otherwise unknown if any child is unknown |
or with any true child | True; otherwise unknown if any child is unknown |
| The whole filter | Selects the item only when true |
The Stripe connector's matches in filter.ts implements these rules, and compares timestamps as nanoseconds. The kit has no filter evaluator, so a connector that re-checks items implements the rules it chooses.
Each Stripe list parameter the connector sends matches exactly, and its created bounds are computed from the exact timestamp. So a single Stripe list request skips the check, while retrieves by id and merged fan-outs go through it. The Art Institute connector translates every filter it accepts into search API queries, so it runs no re-check. It refuses a comparison with the value on the left instead of reversing it.
The check also serves filters Engine composes that the external data source can't express, such as not around an update's field rules and or across permission rules. Where ids bound the candidates, retrieve them and evaluate the rest in memory. Everywhere else, refuse rather than scan.
_null only on a nullable field whose filter declares equals, in a member whose filter also declares not: _null: true arrives as equals with null, and _null: false as not around it. Declare not only if you can answer both, with the external data source's null test or in memory over a bounded set.Page Through Results
Callers page with limit and offset only, and a connector can't declare a maximum for either. See Paging.
Fill the Limit
Engine reads a page shorter than limit as the end of the data. When the external data source's pages are smaller than limit, keep requesting until you reach limit or it runs out, and size the last request to what remains.
| Connector | Page size | limit: 150 costs |
|---|---|---|
| Stripe | 100, capped without an error | A page of 100, then 50 via starting_after |
| Art Institute | 100; size: 101 is refused | from: 0, size: 100, then from: 100, size: 50 |
Keep the Order Stable
Engine adds no sort of its own, so a request without a caller sort arrives without sort. Close every sort with the stable ID, so pages over unchanged data never overlap. Items created, deleted or edited between two page requests can still shift the pages.
The Art Institute connector appends id ascending unless id is already sorted. The Stripe connector breaks created ties by id when it merges requests.
Break ties before anything cuts a result short. When you merge several requests that were each cut at offset + limit, a tie-break applied only after the merge doesn't restore one global order: items that share a sort value at a cut can land on different pages. The Stripe connector avoids this: in a merge, it reads each list on through every item created in the same second as the last one, so every tied candidate reaches the tie-break.
Serve Offset over Cursors
An external data source can offer cursor paging only. Stripe's API reference documents only cursor paging, but Stripe's list endpoints also accept an undocumented offset parameter. The Stripe connector uses it to skip to the caller's offset on the first request, and follows starting_after after that.
On an external data source with cursors only, walk pages and discard items until you reach offset. The cost grows with the page number, so cap the walk.
Refuse Past a Search Window
The Art Institute's search refuses any request that reaches past its 1,000th result. The contract can't declare that ceiling, so the connector checks it before reading any data:
const BASE_URL = 'https://api.artic.edu/api/v1';
/** Engine sets no deadline on `query` or `testConnection`, so every request carries its own. */
const TIMEOUT_MS = 10_000;
/** The search API returns at most 100 items per request... */
const PAGE_SIZE = 100;
/** ...and refuses any request that reaches past its 1,000th result. */
const SEARCH_WINDOW = 1000;
/** A request body for the search API, which takes Elasticsearch syntax. */
type SearchBody = Record<string, unknown>;
// …
/** Answers `readMany`. The response field is `records`; each entry is one item. */
export async function readMany(query: ReadManyOperationSerialized): Promise<QueryResultSerialized> {
const { read, types } = declaration(query.collection);
// Translate everything first, so an undeclared filter or sort is refused before any request.
const base: SearchBody = query.filter === undefined ? {} : { query: toQuery(query.filter, read.filter) };
const sort = toSort(query.sort ?? [], read.sort);
// A count read answers the number of matches and ignores `limit` and `offset`.
if (query.count === true) return { totalCount: await count(query.collection, base) };
const offset = query.offset ?? 0;
let end = query.limit === undefined ? Infinity : offset + query.limit;
// The search API refuses any request past its 1,000th result. A read that reaches past it is
// still safe when fewer items match, so count the matches first and stop at the last one.
if (end > SEARCH_WINDOW) {
end = Math.min(end, await count(query.collection, base));
if (end > SEARCH_WINDOW) {
// A valid read the API can't serve: the caller can change it, so the message says how.
throw new QueryRejected(`The Art Institute search API returns only the first ${SEARCH_WINDOW} results. Narrow the filter or read an earlier page.`);
}
}
const select = query.select ?? [];
const body = { ...base, sort, fields: select.map((name) => apiFieldName(name)) };
const records: MonospaceObject[] = [];
for (let from = offset; from < end; from += PAGE_SIZE) {
const size = Math.min(PAGE_SIZE, end - from);
const page = await search(query.collection, { ...body, from, size });
records.push(...page.data.map((apiItem) => toItem(apiItem, select, types)));
if (page.data.length < size) break;
}
return { records };
}
The connector translates the filter and the sort before the count, so an undeclared sort on an unbounded read is refused before any request. It refuses only a read that needs a result past the window. When the read reaches past result 1,000, it first asks for the total and clamps the end to it:
| Request | Result |
|---|---|
offset=995&limit=5 | Served |
offset=995&limit=10, unfiltered | Refused after one count request |
limit=-1, unfiltered | Refused |
limit=-1 on a filter with 13 matches | Served |
offset=5000 on a filter with 13 matches | An empty page, not a refusal |
Never return a silently shortened page. A caller can't tell it from the end of the data.
Bound Unbounded Reads
Some reads arrive with no limit: a caller's limit=-1 or limit=0 when the instance (your Monospace deployment) sets no maximum page size, and the batched in read Engine sends to resolve a to-one relation, which returns a single related item. A to-many relation, which returns several, arrives as one read per parent with a limit by default. It arrives as one batched read without limit only when the include resolves to no limit and no offset: the include sets limit=-1 or limit=0 on an instance without a maximum, or sets no limit on an instance that unsets its default. See Relation Reads.
Put a budget on what one read may fetch, and refuse when it runs out. The Stripe connector charges each object it takes from a list, and each id it retrieves, against maxItemsPerQuery (5,000 by default, configurable per data source). It refuses once the budget runs out, and an id in filter over the budget is refused before its first request. Its keyed operations retrieve one id and skip the budget.
A budget like this bounds one read, not one caller query. A to-many include that resolves to a limit or an offset, as it does by default, arrives as one read per parent, so a single caller query can spend the budget once per parent. It also doesn't cap network traffic exactly. The connector charges objects as it takes them from a page it has already fetched, so the request that runs the budget out can bring a whole page of up to 100 objects. Concurrent list requests each bring their own page.
Offer Counts Only When the External Data Source Counts
Declare meta.totalCount only if the external data source returns an exact total. A count read arrives as readMany with count: true, limit: 0 and no select. Answer with totalCount for the whole filter, ignoring limit and offset, and return no items. See Count Reads.
The Art Institute connector answers a count read with the same query at size: 0. Stripe's list endpoints return no total, so the Stripe connector doesn't declare it and refuses a count read.
Refuse Reads That Lack a Required Filter
Some lists in an external data source can't be read unfiltered, such as messages that need a channel or items that need a parent. The contract can't declare a required filter, so callers and Studio still offer the unfiltered read.
- Refuse early. Check the filter before any request to the external data source, and throw
QueryRejectedwhen the required constraint is missing. The request is valid by the declaration, and the caller can fix it. - Check every branch. An
orwhose branches don't all carry the constraint is unfiltered too. - Name the fix. Put the missing filter in the message (e.g., "Filter messages by
channelId."). The message is the only explanation the caller sees. - Document it in your connector's README, next to the capability table.
The same applies to bounds the external data source enforces, such as a maximum date range. See Required Filters.
Answer Keyed Operations
readOne, updateOne and deleteOne address one item by its key, plus an optional guard filter. Two implementations work.
Route the key and guard through the read path. When your read path already evaluates filters, combine the key and the guard into one filter. The Art Institute connector reads with limit: 1, after checking the key's shape, and answers the item it finds as record, or record: null:
/**
* Answers `readOne` through the read path: the key, plus the optional guard, at most one item. When
* nothing matches, it answers `record: null`, which callers see as 404.
*/
export async function readOne(query: ReadOneOperationSerialized): Promise<QueryResultFor<'readOne'>> {
const id = query.key.id;
if (Object.keys(query.key).length !== 1 || id === undefined) {
throw new Error(`\`${query.collection}\` keys are a single \`id\`.`);
}
// An id that isn't an integer can't name an item: no item, without asking the API.
if (typeof id !== 'number' || !Number.isSafeInteger(id)) return { record: null };
const key: FilterSerialized = { cmp: 'equals', left: { field: 'id' }, right: { value: id } };
const filter = query.filter === undefined ? key : { and: [key, query.filter] };
const { records = [] } = await readMany({ op: 'readMany', collection: query.collection, select: query.select, filter, limit: 1 });
return { record: records[0] ?? null };
}
Retrieve by key, then check the guard in memory. When the external data source has a get-by-id endpoint, fetch the item and evaluate the guard on it. The Stripe connector's findByKey works this way; see Build a Stripe Data Connector.
Either way:
- A miss answers
record: null. An unknown key, a deleted item and a guard miss all answer{ record: null }, and change nothing. A miss isn't an error, so don't throw. Engine answers the caller with404"No record found". - Check the guard before any write. Read the item, evaluate the guard, then write only what you read.
- Check the key's shape. The key has one entry, the primary-key field. Engine sends only the declared key, so refuse anything else with a plain
Error(e.g., "artworkskeys are a singleid."). - Decide an out-of-domain key locally. Answer it as a miss, without a request. The Art Institute connector answers
record: nullfor an id that isn't a safe integer. The Stripe connector does the same for an id that isn't a string of letters, digits,_and-. - Read an
int64key as a string. It arrives as a string, so never test it withtypeof key === 'number'. - Answer
record, neverrecords. Engine fails a keyed operation that answers{ records: [...] }, even with one item. It checks after your connector has run, so a write it made stays applied. Type the function's answer asQueryResultFor<'readOne'>(or the operation you serve), andtscrefusesrecords. - Treat tombstones as missing. Stripe still returns a deleted customer, with
deleted: true, and the Stripe connector treats it as a miss.
Write Safely
- Declare what your writes return. When your connector answers each kind of write with the items it wrote, carrying every field in
select, declare the matching query capability. The capabilities apply to every collection the entry serves. The Stripe connector declares all three. Without a capability, Engine makes extra calls by stable ID: it askscreatefor the stable IDs and reads the created items back, and it sends updates and deletes throughupdateManyanddeleteMany, betweenreadManycalls. That needs a stable ID and declarations that carry those calls. See Declare Query Capabilities. - Return the selected fields. Answer every write with the items it wrote, carrying every field in
select, with or without a capability. The capabilities change what Engine selects: acreatewithoutinsertReturningselects only the stable ID, and a write withselect: []can answer with norecords. When the external data source returns nothing useful (Stripe answers a delete with a tombstone), read the selected fields first, then write. - Write by id when the external data source can't write by filter. Resolve the filter with your own read, then write each id. The Stripe connector does this for
updateManyanddeleteMany. - Accept that writes aren't atomic. Engine runs no transaction across connector writes. A failure partway leaves earlier writes applied, and a concurrent writer can change an item between your read and your write. Document it.
- Treat a failed request's outcome as unknown. A write that times out, or whose answer can't be read, can still have been applied. A count of acknowledged writes is a lower bound, so reconcile by reading, or use idempotency keys where the external data source offers them.
- Stop at the first failure, and say how far you got. In the Stripe connector's fan-outs, no new write starts after a failure; the writes already in flight finish, then the error is thrown. For a write of several items, the message says how many were written and keeps the error's class: "Deleted 3 of 4 products before this error: This product cannot be deleted because it has one or more user-created prices."
- Bound concurrency. The Stripe connector runs each fan-out with at most four concurrent requests. That caps concurrency, not requests per second, and the cap is per fan-out, so nested fan-outs and concurrent queries add up. Staying under a requests-per-second limit needs pacing of its own.
- Map refusals from the external data source to constraint errors. Stripe refuses to delete a product that has prices. The Stripe connector reports it as
ForeignKeyConstraintViolation.
Defaults protect data the caller didn't mean to touch:
- Default writes off for credentials that reach production data. The Stripe connector's
enableWritesdefaults to on only for test-mode keys. With writes off, it declares a read-only schema, so Engine refuses writes before they reach the connector. - Give side effects their own switch. When creating an item sends something (an email, an SMS, a charge), gate it behind a configuration option separate from
enableWrites. - Don't repeat side effects on retry. If the external data source supports idempotency keys, send one per item and reuse it when you retry after a network error or a
5xx.
The Stripe example keeps its request path small: it sends no idempotency keys and doesn't retry. Every request has a 10-second deadline. Add idempotency keys and retries where your external data source needs them.
Map Values, Types and Keys
Declare a Static Schema
When the external data source has no metadata API, declare the schema in code.
- Declare each collection statically. Write the
operationswithfields.all()for every field, orfieldsOf(fields).pick(...)for a list several members share, sotsccatches a typo. A field is writable exactly when adatalist names it; the kit derives the read-only flags. The Stripe connector'scollections.tsdoes this. - Keep the mapping next to the declaration. Key it by the declared field names, so a typo there is a type error too. The Stripe connector's
resources.tsholds each field's path in the Stripe object, its form key and its list parameter, and reads what the field can do from the declaration. - Pin the API version in every request, and change the pinned version and the field paths together.
- Check for drift. Test that every declared field reads a value at its path on a live object.
Map Types from Metadata
When the external data source describes its schema, derive the collections from it at introspection.
- Map each type the external data source reports to a field type, and degrade honestly. A decimal becomes
decimal, returned as a string:field.decimal({ precision, scale })when the external data source declares both, andfield.decimal()when it declares neither. An unknown type becomesunsupported, which no filter,outputordatalist can name, so callers can't read it. See Match Operators to Field Types. - Check value ranges before you pick a type. The Art Institute's
date_startisn't limited to plausible years: it holds values such as-1824528578, which still fitsint32. The connector declares its ids and years asint32, so they travel as JSON numbers in data, filters, keys and REST responses. Engine sendsint64anduint64values as decimal strings in item data, filter values and keys, accepts a string or a number back, and returns them to callers as strings. Return values outside JavaScript's safe-integer range as decimal strings. That keeps answers exact only if you read them exactly:JSON.parsehas already rounded such a number, so parse the external data source's response in a way that keeps the digits, or fail on the value. - Derive nullability from data where the metadata is silent. Count items where the field is missing. A field with missing values must be nullable, but finding none doesn't prove future values are present. Declare a field non-nullable only where the external data source guarantees a value.
- Declare only keys you can prove. Declare a primary or unique index only where the external data source guarantees it. Keyed operations need a single-field primary key. When the external data source identifies an item by several values, expose one string field that combines them as the primary key, and split it again in the connector; an index over several fields doesn't enable keyed operations.
- Declare relations with key fields of the same type. Engine offers each direction only when the target collection can filter by the key and its members list the key fields in
output. For a single key field, the filter needsinandand; withoutin,equalswithandandoralso works, and Engine then sends oneequalsper key value joined withor. For a key of several fields, it needsequalson each field, withandandor. So one direction can stay available when the other can't. A key whose type can't takein(e.g., a singlebooleanfield), or that contains anunsupportedfield, is exempt from the filter requirement, and Engine sends its relation reads without a declaration; see Rules for Engine's Own Calls and Resolve Relations.
Normalize Values
Return every value in the encoding its declared type expects, and project exactly the fields in select. Carry every selected field: use an explicit null for a missing value, because a missing field fails the request. Check each value against its field's type and nullability as you convert it, so a changed API fails in your connector instead of with a type error in Engine. Engine doesn't check nullability at all: it accepts null for a non-nullable field. Throw a plain Error for it: the caller can't fix a changed API, and Engine reports it as a failed query. The Stripe connector fails on a missing value for a field that isn't nullable, and converts Unix-second timestamps to RFC 3339, failing on a number too large for a date:
import type { MonospaceObject, Value } from '@monospace/extension-kit/data-connector';
import type { StripeObject } from './client';
import type { FieldSpec, ResourceSpec } from './resources';
import { QueryRejected } from '@monospace/extension-kit/data-connector';
import { isObject } from './client';
import { apiFieldName } from './resources';
/** An item, with the Stripe id the write code addresses its object by. */
export type Item = MonospaceObject & { readonly id: string };
/** Converts a Stripe object to an item with every declared field, each checked against its kind. */
export function toItem(resource: ResourceSpec, object: StripeObject): Item {
const item: MonospaceObject = {};
for (const field of resource.fields) item[field.name] = toValue(field, readPath(object, field.path ?? [apiFieldName(field.name)]));
return { ...item, id: object.id };
}
// …
function readPath(object: StripeObject, path: readonly string[]): unknown {
let value: unknown = object;
for (const key of path) {
if (!isObject(value)) return undefined;
value = value[key];
}
return value;
}
function toValue(field: FieldSpec, raw: unknown): Value {
if (raw === undefined || raw === null) {
if (field.isNullable) return null;
throw new Error(`Stripe returned no value for \`${field.name}\`.`);
}
switch (field.kind) {
case 'id':
case 'string':
if (typeof raw === 'string') return raw;
break;
case 'int64':
if (typeof raw === 'number' && Number.isSafeInteger(raw)) return raw;
break;
case 'boolean':
if (typeof raw === 'boolean') return raw;
break;
case 'timestamp': {
// Stripe sends Unix seconds. Engine reads RFC 3339. A number outside the range `Date` can
// hold makes an invalid date, whose `toISOString` would throw a plain `RangeError`.
const date = typeof raw === 'number' ? new Date(raw * 1000) : undefined;
if (date !== undefined && !Number.isNaN(date.getTime())) return date.toISOString();
break;
}
}
// A changed API, not something the caller can fix: a plain error, which Engine reports as a failed query.
throw new Error(`Stripe returned an unexpected value for \`${field.name}\`.`);
}
See Field Types for each type's encoding.
Map Errors Precisely
The error class you throw decides what the caller sees. See Handle Data Connector Errors for the full mapping and Data Connector Errors for each class.
- Only a vendor-shaped "not found" is a miss. Classify by the external data source's error body, not the status alone. A bare
404from a wrong base URL stays a plainError, or a misconfigured data source reads as empty. - Handle a miss where it means no match. The Stripe connector maps
resource_missingby its error code. It readsparamto tell a missing object from a missing reference. Itsretrieveanswers a404with that code as no object: a keyed read answersrecord: null, and a read by id leaves the item out. A list filtered by a missing referenced object answers no items. - A missing object elsewhere is a plain
Error. No kit class reports a missing item. Aresource_missingerror that reaches the Stripe client's classifier becomesStripeObjectMissing, the connector's own subclass ofError. Its keyed writes answer it asrecord: null. In a bulk write or on a later page, callers see a failed query. - A missing reference on a write is a constraint violation. When a write names an item that doesn't exist, throw
ForeignKeyConstraintViolation. The caller then gets422instead of a failed query. Stripe answers a price create with a missing product asresource_missingwithparam: product, and the Stripe connector maps it that way. - A rejected credential is
AccessDenied. Only someone with access to the external data source's account can fix it. The Stripe connector maps401and403toAccessDenied, so a wrong key fails the connection test with a clear message.InvalidConfigurationis for a value that's wrong before it's ever sent. - Put the retry hint in the message. Besides the error's class, the message is all the caller gets. Both example connectors throw
RateLimitedfor a429, and append "Retry after 2 seconds." whenRetry-Afteris a whole number of seconds, ignoring an HTTP-date. Nothing retries on it, so retry inside the connector where the request is safe to repeat. - Server errors and timeouts are
ConnectionFailed. A5xxmeans the external data source failed to serve this request, and a timeout means it didn't answer in time. Both examples treat either as an outage: a later request may work, so a caller can retry where the request is safe to repeat. Any other failed status the connector has no rule for is a plainError. testConnectionauthenticates. Make it an authenticated read, so a wrong credential fails the test instead of the first query.- Keep secrets out of messages. The message is the only text of an error that reaches the caller. Stripe's error bodies include a URL that contains the account id, and its message for a rejected key echoes part of the key. The Stripe connector uses only Stripe's
messagefield, and a fixed message for a rejected key. The Art Institute connector puts no response body into a message. - Give every request a deadline. Engine sets none. Both example connectors pass
AbortSignal.timeout(10_000)tofetchand read the body inside the sametry.
console output from your connector lands in the Engine log at info level, tagged with the data source and the connector. Log the request you sent to the external data source when a mapping surprises you.
See Also
- Build a Stripe Data Connector: the Stripe connector, step by step
- Operations and Operators: what a collection can declare, and the rules Engine enforces
- Data Connector API: the query wire format and the shapes Engine sends on its own
- Data Connector Limitations: what the contract can't express
- Test a Data Connector: check every mapping against the external data source