Build Data Connectors with AI Agents
@monospace/cli 0.2 and @monospace/extension-kit 0.3. Overview
A coding agent can build an extension end to end. Point it at these docs instead of its memory, and stay in the loop for the steps it can't check. Each entry type has its own workflow.
Data connectors are the only entry type today, so this page covers building a data connector extension. It exposes an external data source (the API, service, or database the connector reads and writes) as collections in a Monospace workspace. The agent probes the external data source, writes the connector, and climbs the test ladder.
This page covers what to give the agent, a prompt to start from, the rules to hold it to, and what an agent can't verify alone.
Give the Agent the Docs
Every docs page has a markdown version, and the site publishes two index files for language models.
| URL | Contents |
|---|---|
https://docs.monospace.io/llms.txt | Every page's title, description, and markdown URL |
https://docs.monospace.io/llms-full.txt | The full text of every page in one file |
https://docs.monospace.io/raw/en/<page path>.md | One page as markdown (e.g., https://docs.monospace.io/raw/en/developer/extensions/data-connectors/runtime.md) |
https://docs.monospace.io/en/<page path> | One page as HTML (e.g., https://docs.monospace.io/en/developer/extensions/data-connectors/runtime) |
Give the agent the markdown URLs of the extension pages rather than llms-full.txt. The full file covers the whole platform, and most of it is unrelated to extensions.
Read these pages in this order. The agent needs the runtime model and the limitations before it writes code:
| Page | Why the agent needs it |
|---|---|
| How Data Connectors Run | The sandbox: one file, no Node APIs, configuration instead of environment variables |
| Data Connectors | Live queries, the operations contract, and the Fit Check |
| Limitations | What a connector can't do, and the workaround for each |
| Build a Stripe Data Connector | The build, step by step |
| Map an External Data Source | Probing, filters, paging, keys, and writes |
| Handle Errors | Which error class to throw for each failure of the external data source |
| Test a Data Connector | The test ladder |
| Data Connector API, Operations and Operators, Errors, Network Permissions, CLI | Exact shapes, rules, and codes to check against |
| Data Connector Module, Protocol Types, Schema Helpers | The declaration of every export of @monospace/extension-kit |
Start with a Prompt
Copy this prompt, fill in the angle-bracket parts, and give it to your agent in the directory where the connector project goes. It points the agent at the docs, sets the safety defaults, and makes it stop at the checkpoints where you decide.
Build a Monospace data connector extension for <EXTERNAL DATA SOURCE, e.g. "the Acme REST API at https://api.acme.example">.
Expose <COLLECTIONS, e.g. "customers and invoices, with invoices -> customers"> as collections.
Writes: <"none, read-only" | "create and update on customers, behind an enableWrites option">.
Read these pages first, and trust them over anything you remember about Monospace. Re-read the
relevant page before each step:
- https://docs.monospace.io/raw/en/developer/extensions/data-connectors/runtime.md
- https://docs.monospace.io/raw/en/developer/extensions/data-connectors.md
- https://docs.monospace.io/raw/en/reference/extensions/data-connectors/limitations.md
- https://docs.monospace.io/raw/en/guides/data-connectors/stripe.md
- https://docs.monospace.io/raw/en/developer/extensions/data-connectors/data-source-mapping.md
- https://docs.monospace.io/raw/en/developer/extensions/data-connectors/errors.md
- https://docs.monospace.io/raw/en/developer/extensions/data-connectors/testing.md
- https://docs.monospace.io/raw/en/reference/extensions/data-connectors/api.md
- https://docs.monospace.io/raw/en/reference/extensions/data-connectors/operations.md
- https://docs.monospace.io/raw/en/reference/extensions/data-connectors/permissions.md
- https://docs.monospace.io/raw/en/reference/extensions/data-connectors/errors.md
- https://docs.monospace.io/raw/en/reference/extensions/cli.md
- https://docs.monospace.io/raw/en/reference/extensions/extension-kit/data-connector.md
- https://docs.monospace.io/raw/en/reference/extensions/extension-kit/protocol.md
- https://docs.monospace.io/raw/en/reference/extensions/extension-kit/experimental.md
The index of all pages is https://docs.monospace.io/llms.txt.
Rules:
- Run the Fit Check before building. If a limitation blocks what I asked for, stop and ask me.
- Probe the live external data source before declaring anything. Every filter, sort, count and
paging option you declare must name a request to it that you ran. Don't trust the vendor's
docs alone.
- Declare only what requests to the external data source can answer, and re-check every returned
item against the filter where its match is broader. Refuse the rest before any request to the
external data source when the request's shape decides it: QueryRejected for a valid request it
can't serve (including shapes Monospace composes: not, or), a plain Error for anything the
collection doesn't declare, anywhere in the filter. Monospace doesn't check the filters it
sends, so a permission rule's condition or a stale stored declaration can bring an undeclared
comparison: refuse it, never drop it. Exception: serve the relation and read-back
lookups Monospace sends undeclared on a key no declaration can carry (in on a one-field
boolean, bytes, json, geometry or unsupported key; equals on each field of a key with an
unsupported field). Never scan, and never truncate a page.
- Bound every read. Monospace sends some reads without a limit (limit=-1, relation batches), so
refuse with QueryRejected once a read would fetch more than a budget you set.
- Declare collections with the kit's builders: fields.all(), fieldsOf(...).pick(...),
paging.limit() and paging.offset(). A field is writable exactly when a data list names it.
- Declare a query capability (insertReturning, updateReturning, deleteReturning, unique values
in the entry's engine.capabilities.query, applying to every collection of the entry) only for
writes whose answers carry the written items with every selected field. Updates always need
readMany to read their items back by stable ID. A collection with a write whose capability is
missing needs a primary or unique index. Without insertReturning, create.output lists the
stable-ID fields, create answers them, and readMany must read the created items by stable
ID. Without updateReturning, updates go through updateMany by stable ID, and no update's
data names a stable-ID field. Without deleteReturning, deletes go through deleteMany by
stable ID, after readMany reads the items by stable ID.
- A keyed operation (readOne, updateOne, deleteOne) answers { record: <item> }, or
{ record: null } when its key or guard misses, and then writes nothing. Never throw for a
miss, and never answer records. Type it with QueryResultFor<'readOne'> (or its operation).
int64 values and keys arrive as strings.
- Fixed hosts go in extension.config.json as net: permissions. Use preflight only for hosts
derived from configuration. Validate configuration in setup and throw InvalidConfiguration;
preflight may also throw InvalidConfiguration when the configuration it reads is invalid.
- No Node APIs, no environment variables, no filesystem. Credentials come from the data source
configuration.
- Use test-mode or sandbox credentials only. Default writes off for credentials that can reach
production data.
- Keep secrets and personal data out of error messages: the message is the only text of an
error that can reach the caller. Never copy a response body or a query string into one.
- Give every request to the external data source a timeout (AbortSignal.timeout), and read the
body inside the same try. Monospace sets no deadline on query, introspect or testConnection.
- Map failures in one client: no answer, a timeout or a 5xx is ConnectionFailed, a 429 is
RateLimited, a rejected credential (401/403) is AccessDenied. A missing required value, a
duplicate, or a missing reference is NullConstraintViolation, UniqueConstraintViolation or
ForeignKeyConstraintViolation; any other value the external data source refuses is
QueryRejected. Wrap a fetch rejection with { cause: error }, so a permission denial
stays one. Throw a plain Error only for a changed API (a body that isn't JSON, a wrong shape)
and for a request the declaration never offered. Compare error codes with ErrorCode, never
with hard-coded numbers.
- Check every value before you return it: every selected field present, each value of its
declared type, null only on nullable fields, and integers within their type's range.
- Writes aren't atomic. Read a keyed write's guard before writing, report how far a bulk write
got when it fails, and never retry a write that isn't safe to repeat.
- If the external data source takes a credential, implement testConnection and test it with a
malformed, a wrong, and a valid credential. If you don't implement it, say so in your report.
Work in this order, and stop to report at each checkpoint:
1. Fit Check, probes, and a capability table (what you'll declare, and the request to the external
data source that answers each row). CHECKPOINT: show me the table and the limitations that apply.
2. Project, configuration parser, client with error mapping, and the schema. Until reads exist,
make query throw QueryRejected. Typecheck, build, install, and request a data source
in a throwaway workspace. CHECKPOINT: when Monospace answers 7001, show me the required network
permissions before you resend them. Then create the data source before writing query code.
3. Reads, then keyed operations, then writes, each with network-free unit tests (a fake fetch for
timeouts, 429s and bodies that aren't JSON) and artifact smoke checks: wire assertions for
every translated filter, no-request assertions for every shape-based refusal.
4. The Engine smoke, one deliberate break per check, and a second full run. Send every Engine read
with Cache-Control: no-cache, so a cached result can't hide a connector failure.
5. A README with the configuration, network access, the capability table, refusals, and gaps.
When you finish, list every test rung you ran, every rung you skipped, and everything you could
not verify yourself.
The checkpoints matter more than the rules. The first one catches an external data source that doesn't fit before any code exists, and the second keeps network approval with you.
Rules for the Agent
Hold the agent to these rules. Each links to the page that owns the detail.
| Rule | What goes wrong without it |
|---|---|
| Run the Fit Check first, and ask when a limitation blocks the request. See Fit Check | The agent builds around a required filter or an unwritable reference, and the user finds out in Studio |
| Probe the live external data source, and declare only what a request you ran answers. See Map an External Data Source | Declarations based on vendor docs fail at runtime. The Art Institute of Chicago API documents 10,000 search results, and refuses past 1,000 |
Put fixed hosts in the manifest. Use preflight only for hosts derived from configuration. Validate the configuration in setup. preflight may throw InvalidConfiguration for the part it reads. See Network Permissions | A preflight that declares fixed hosts hides the grant from the manifest |
| Write web-platform code only. See How Data Connectors Run. Need a Node.js-compatible API? Let us know | node:* imports fail the build, and Deno.env or filesystem calls throw NotCapable at runtime |
Run the linter, tsc and unit tests; the build doesn't typecheck. See Lint, Typecheck, and Run Unit Tests | A green build ships a type error, or an error mapping that only a timeout or a 429 would exercise |
| Create a data source before writing query code, and check every relation in the workspace's OpenAPI document. See Test a Data Connector | Engine checks declarations when a data source is created or reintrospected, not at build time or on the raw /sources/data/introspect endpoint. A reload doesn't refresh a data source's stored declarations either; reintrospect after changing them. It narrows a relation direction it can't serve with only a log warning |
Serve or refuse: never scan, never truncate. Refuse with QueryRejected (or a plain Error for what you didn't declare) as early as the connector can tell, fill the whole limit unless the matching items run out, and break sort ties with the stable ID. See Handle Errors | A short page reads as "no more items", and a scan burns the external data source's rate limit |
Re-check every returned item against the filter when the external data source's match is broader than the declared operator (e.g., case-insensitive or substring), and keep fetching when the re-check drops items, so the page still fills its limit. See Map an External Data Source | Callers get items that don't match their filter, or a short page that reads as the end of the data |
Answer { record: null } for a keyed miss, and check the guard before any write. See Keyed Operations | A guard miss still writes, or answers an item the guard excludes. A plain Error or a records answer fails the caller's request with 500 instead of 404 |
Map errors in one client, use InvalidConfiguration only for configuration and AccessDenied only for an account the external data source refuses (rejected credentials, a missing scope, or a resource it can't reach). Give every request a deadline, and wrap a fetch rejection with { cause }, so a permission denial stays one. See Handle Errors | Callers are told their configuration is invalid when the external data source is down, see an internal error, or wait on a request that never answers |
| Assert what goes on the wire, and make every check fail once. See Test a Data Connector | Checks pass against a fixture that ignores the parameter being tested |
| Report honestly. List skipped rungs, and exit non-zero from a run that skipped checks | "All tests pass" means only the rungs that ran |
Know What the Agent Can't Verify Alone
Some steps need a person, credentials, or an instance (a Monospace deployment) the agent can't reach. Plan for these before you start, and ask the agent to list them in its final report.
| Step | Why the agent needs you |
|---|---|
| Approving network permissions | Deciding which hosts a data source may reach is your call. Review meta.permissions before the agent resends them. The grant must contain exactly the required hosts, ports and wildcards; reasons aren't compared |
| Credentials | The agent needs a test-mode or sandbox credential from you. Never give it production credentials |
| Installing into your instance | The agent needs write access to the instance's install location, and a restart or automatic reloading. See Install and Manage Extensions |
| Behavior that needs live traffic | Real rate limits and paging depth show up only against the real external data source, often only under load. The agent can test its 429 and 5xx mapping alone with a stand-in that injects them |
| Studio | Browsing, filtering, and editing in Studio, and judging whether each error message helps a user. No example automates it. See Try It in Studio |
See Also
- Build a Stripe Data Connector: the Stripe connector, step by step
- Test a Data Connector: the ladder an agent climbs before it reports done
- Limitations: what a connector can't do, for the Fit Check
- How Data Connectors Run: the sandbox the agent's code runs in