# Import, export and external publishing

**For site owners and integrators.** Kartotek's **native JSON interchange format** for a Post — a lossless, versioned document that can be downloaded from a Published Post's static export, uploaded through the admin editor, or submitted by an external system through a scoped API. See [Welcome to Kartotek](/_docs/kartotek) for the vocabulary this doc assumes (Post, Version, ContentElement, Role, Rank, Reference).

## Getting started

**Two export formats, two different jobs:**

- **JSON-LD** (`/<uid>/metadata/current` and friends — see [Posts and publishing](/_docs/posts)) — descriptive metadata for a generic consumer (search engines, link previews, anything that understands Schema.org). Not lossless — don't use it to move a Post between Kartotek instances.
- **Native JSON** (this doc) — a complete, lossless, re-importable document of a Post's own structure. Use this to back up a Post, move it to a different uid, or publish into a Kartotek instance from an external system.

**Public backup vs. self-contained admin export:** every Published Post has a public export at `/_static/<uid>.json` (media referenced by URL, Hidden Versions excluded). An admin can additionally download a self-contained variant (media embedded as base64, Hidden Versions included) for a Post in any Stage — see "Downloading an export," below.

**Simplest way to import:** open the admin sidebar's **Imports** section, paste or upload a `.json` document, and submit — see "Admin upload walkthrough," below.

**Simplest way to publish externally**, one request with a scoped token:

```bash
curl -X POST https://hoejriis.dk/_api/import/posts \
  -H "Authorization: Bearer kit_xxxxxxxx…" \
  -H "Content-Type: application/json" \
  -d '{"document": {"format":"kartotek-post","formatVersion":2,"post":{"uid":"external-post"},"contentElements":[],"versions":[]}, "mediaMode": "automatic", "autoFinalize": true}'
```

A `201` means it finalized immediately into a new PrePub Post; anything else stages the job for review — see "Integration reference," below, for the full endpoint set and response codes.

**Archive Media from an external producer** uses its own format and endpoints: see "Media import (Archive Media)," below.

## Downloading an export

Every **Published** Post has a static, public export alongside its PDF (see [Posts and publishing](/_docs/posts)'s "Static PDF export"):

**`GET /_static/<uid>.json`**

- native `kartotek-post` format, current version 2;
- the Post's complete **public** Version history (Hidden Versions excluded entirely — not included and marked hidden, simply absent);
- media referenced by absolute URL, never embedded bytes;
- regenerated on every publish while the Post stays Published, deleted the moment it leaves Published — same lifecycle as the PDF.

An admin can also download a **self-contained** export — Hidden Versions included, media embedded as base64 data URIs — for a Post in any Stage (including PrePub), from the admin editor or directly:

**`GET /_api/admin/posts/<uid>/export.json?includeHidden=true&embed=true`** (admin session required)

Both query parameters default to `false`; omit either to get the public-equivalent shape from an authenticated request.

## Admin upload walkthrough

From the admin sidebar's **Imports** section: paste a document or choose a `.json` file, pick a media mode (**Automatic** — embedded data first, otherwise fetch the URL; **Embedded only**; **Fetch URLs**), and submit. A clean import with every media item acquired successfully finalizes immediately. Anything else stages the job for review: failed media items can be repaired individually — replace the source URL or upload a replacement file directly — then retried, or the whole job finalized anyway with the still-failing items explicitly omitted (**Finalize without failed media**, with a confirmation listing exactly what's being dropped). A staged job can also be abandoned outright, which deletes any already-acquired media files and frees the uid.

## Integration reference

The rest of this doc is detailed reference material — the exact document schema, every external-API endpoint and response code, security constraints, and troubleshooting — for building or debugging an integration, not needed to just download a backup or use the admin import UI above.

### Troubleshooting, by symptom

- **`409 Conflict` on submit** — the requested uid is already taken (by a Post, a Category, a historical alias, or another import already in progress for it). Change `post.uid` and resubmit.
- **Post imports uncategorized unexpectedly** — `post.categoryUid` didn't match an existing Category; check the import's report for the warning, and re-assign the Category manually after import if needed.
- **A media item stays failed after retrying** — the source URL may be unreachable, return an unexpected content-type, or exceed the size limit (see "Security," below); replace it with a working URL or upload the file directly instead.
- **`400 Bad Request` with no job created** — the document itself is structurally invalid (bad format/version, invalid Role, a Version referencing an unknown ContentElement key). Check the error detail against "Document structure," below.
- **A document is rejected for its `formatVersion`** — this deployment only understands the versions listed in "Format versioning," below; re-export from a compatible version or wait for this deployment to support a newer one. An *older* supported version (currently: 1) is never rejected for that alone — see "Format versioning."

### Document structure

```json
{
  "format": "kartotek-post",
  "formatVersion": 2,
  "exportedAt": "2026-07-19T06:00:00.000Z",
  "source": { "baseUrl": "https://hoejriis.dk", "uid": "example-post" },
  "post": { "uid": "example-post", "categoryUid": "photos", "stage": "published", "sourceLanguage": "en" },
  "contentElements": [],
  "versions": []
}
```

| Field | Required | Meaning |
|---|---:|---|
| `format` | yes | Must equal `"kartotek-post"`. |
| `formatVersion` | yes | Integer; this deployment currently understands `1` through `5` (see "Format versioning," below). An unrecognized version is rejected, not guessed at. |
| `exportedAt` | no | Informational timestamp; ignored on import. |
| `source.baseUrl` / `source.uid` | no | Informational — where this export came from. |
| `post.uid` | yes | The uid a fresh import will use. Editable in the admin upload flow before starting. |
| `post.categoryUid` | no | An existing Category's uid. Unresolved on import → a warning, imported uncategorized. |
| `post.stage` | no | Informational only — an import always starts PrePub regardless of this value. |
| `post.sourceLanguage` | no | The Post's one authoritative language — see [Posts and publishing](/_docs/posts)'s "Translation" section. An unrecognized value → a warning, imported using the site default instead. Never treated as authored translated content — an importer declares this as fact, it isn't detected. |
| `contentElements` | yes | The Post's deduplicated ContentElement pool (see below). |
| `versions` | yes | The complete, ordered Version history, each referencing the pool by key. |

Database ids are never part of this format. Every ContentElement instead gets an export-local key, `<kind>:<n>` (e.g. `image:1`, `text:3`), assigned in first-appearance order across the Version history. A ContentElement shared by several Versions (carried forward unchanged, per [Posts and publishing](/_docs/posts)'s structural-sharing rule) appears **once** in `contentElements` and is referenced by that same key from every Version that uses it — reimporting reproduces that sharing rather than duplicating rows.

### ContentElement shape

```json
{
  "key": "image:1",
  "kind": "image",
  "role": "cover",
  "rank": 1,
  "reference": "hero-image",
  "createdAt": "2026-07-19T05:00:00.000Z",
  "contentKey": "vLyOl-R21giF",
  "data": {}
}
```

`contentKey` is this element's own stable, permanent identity — distinct from `key` above, which is only a within-this-document cross-reference token, regenerated fresh on every export. `contentKey` survives round-tripping: re-importing an exported document restores each element's original identity instead of minting a new one. Optional on import (an export produced before this field existed simply omits it, and a fresh identity is assigned automatically); when present, it's carried through as-is, not validated against any particular format.

`kind` is one of `image`, `text`, `link`, `contact`, `video`, `geospatial`, `timeline`, `audio`, `attachment`. `role`/`reference` follow the same validation as everywhere else in Kartotek (lowercase letters, numbers, hyphens — see [Posts and publishing](/_docs/posts)). `data`'s shape depends on `kind`:

| Kind | `data` fields |
|---|---|
| `text` | `bodyMarkdown` (string), `format` (`"markdown"` \| `"html"`, optional — see "Text format and the shared metadata envelope," below), `metadataEnvelope` (optional) |
| `link` | `url`, `label`, `note`, `metadataEnvelope` (optional) |
| `contact` | `contactKind` (`email`\|`phone`\|`social`), `value` |
| `geospatial` | `geometryType` (`point`\|`linestring`\|`polygon`\|`heatmap`), `points` (array of `{lat, lng}` — every Geometry Type is a points list, see the Geospatial Content Type's schema page), `label`, `metadataEnvelope` (optional) |
| `timeline` | `title` (optional), `precision` (`"year"` \| `"date"` \| `"time"` \| `"date_time"` \| `"full"`), `events` (array of `{timestamp, label, description, rank}` — see the Timeline Content Type's schema page). A `formatVersion` 1-3 document still uses the older single-date shape (`dateStart`, `dateEnd`, `label`, `description`) — see "Format versioning," below. |
| `image` / `video` / `audio` / `attachment` | see below — `metadataEnvelope` (optional) is additionally supported on `image`, `video`, and `audio`; `attachment` additionally supports `searchText` (optional); all four additionally support `categoryUids`/`tags` (optional — see below) |

#### Image, Video, Audio, Attachment

Embed form (no local file — see [Media and the Media Archive](/_docs/media); Attachment never uses this form, since it has no Embed mode):

```json
{ "mode": "embed", "embedHtml": "<iframe …></iframe>", "caption": "…" }
```

File form — url-only (the static public export's shape):

```json
{ "mode": "file", "contentUrl": "https://hoejriis.dk/_media/123.jpg", "embeddedData": null, "caption": "…", "metadata": null }
```

File form — self-contained (the admin export's shape, `embed=true`):

```json
{ "mode": "file", "contentUrl": "https://hoejriis.dk/_media/123.jpg", "embeddedData": "data:image/jpeg;base64,…", "caption": "…", "metadata": null }
```

`contentUrl` is kept even when `embeddedData` is present, as provenance. `metadata` mirrors Kartotek's own extracted EXIF/capture data where present, but an importer always re-extracts fresh metadata from the actual bytes rather than trusting this field. For an `attachment`, `data` additionally carries `searchText` — the PDF/plain-text content extracted at the original ingestion (see [Media and the Media Archive](/_docs/media)) — which an importer carries through verbatim rather than re-extracting.

**`categoryUids` and `tags`** (both optional, on `image`/`video`/`audio`/`attachment` only) preserve that element's own Category/Tag relationships (see [Organizing and presenting content](/_docs/categories)'s "Categories and Tags on media") across export/re-import — omitted entirely when the element has neither, the same "omit rather than export an empty/null placeholder" rule `post.categoryUid` already follows. `categoryUids` is an array of existing Category uids; `tags` is an array of Tag spellings, same as a Version's own `tags` field. An unresolved `categoryUid` (a Category renamed or detached elsewhere) or a malformed Tag is skipped individually on import rather than failing the whole document — this is best-effort relationship carry-forward, not a required field. Re-importing under a new `post.uid` (rather than updating the original Post) never affects the source element's own relationships, only the newly-imported one's.

### Text format and the shared metadata envelope

Text supports two storage formats, declared explicitly per element rather than guessed at: `"format": "markdown"` (the default when omitted — an older `formatVersion 1` export is implicitly this) or `"format": "html"`, a trusted, explicit escape hatch for legacy content that doesn't convert cleanly to Markdown. A Text element's `bodyMarkdown` field holds whichever format its own `format` says, never a mix.

Text, Image, Link, Geospatial, Video, and Audio ContentElements (not Contact/Attachment/Timeline) may additionally carry a `metadataEnvelope` object:

```json
{
  "label": "Original caption",
  "description": "Longer free-text description",
  "altText": "Alt text (Image only, in practice)",
  "creator": "Attribution / byline / credit",
  "sourceLabel": "Human-readable source name",
  "sourceUrl": "https://original-source.example/item/123",
  "originalDate": "1998-06-14T00:00:00.000Z",
  "location": { "lat": 55.6, "lng": 12.5 },
  "captureAngle": 180,
  "rights": "CC-BY-4.0",
  "language": "da",
  "fieldOrigin": { "label": "imported", "description": "owner" },
  "provenance": {
    "sourceSystem": "legacy-cms-name",
    "sourceItemId": "12345",
    "converterName": "legacy2kartotek",
    "converterVersion": "1.0.0",
    "importedAt": "2026-07-23T00:00:00.000Z"
  },
  "extensions": {
    "legacycms": { "legacyId": "12345", "legacyStatus": "archived" }
  }
}
```

Every field is optional, and import is lenient: a malformed individual field (an invalid `originalDate`, out-of-range `location`, an unrecognized `fieldOrigin` value) is dropped with a warning rather than failing the whole element or document. `extensions` is a namespaced escape hatch for source-specific fields with no first-class home above — it's inert (never read by any rendering code) and passes through a safety filter: a key name suggesting a credential/session/token, a value that looks like serialized legacy code, active markup, an IP address, or an internal filesystem path is dropped (with a warning); nested objects/arrays aren't supported, only flat scalars; and the whole `extensions` object is dropped past a total size bound. This filtering is best-effort defense-in-depth, not a substitute for an external converter only ever producing safe values in the first place.

### Version shape

```json
{
  "ordinal": 2,
  "createdAt": "2026-07-19T05:30:00.000Z",
  "postType": "article",
  "title": "Example Post",
  "hidden": false,
  "hiddenReason": null,
  "content": ["image:1", "text:2"],
  "tags": ["woodworking", "journal"],
  "translations": [
    { "language": "da-DK", "status": "succeeded" },
    { "language": "en", "status": "original" }
  ],
  "sourceProvenance": null
}
```

`ordinal` is the Version's position in the Post's *original* full history (see [Posts and publishing](/_docs/posts)'s "the Nth-ever-published Version") — strictly ascending, but a **public** export (Hidden Versions excluded) can have gaps where a Hidden Version used to be. A gap is a warning on import, not an error: the missing ordinals are reported, and the imported Post's own local ordinals come out contiguous (no fabricated Versions are ever created to fill a gap). `content` is the ordered list of ContentElement keys this Version references — normal Post Types are supported (`article`, `link`, `image`, `gallery`, `category_presentation`, `fullscreen`, `facebook_post`, `instagram_post`, `linkedin_post`), each subject to the same Display Schema described in [Posts and publishing](/_docs/posts) and the three social Types' own Schema pages ([Facebook Post](/_docs/facebook_post), [Instagram Post](/_docs/instagram_post), [LinkedIn Post](/_docs/linkedin_post)). `tags` is this Version's own Tags (see [Organizing and presenting content](/_docs/categories)'s "Tags" section) in authored display spelling — optional, omitted or empty means no Tags; an individual malformed entry is dropped with a warning rather than failing the whole document, since there's no registry for a Tag to be valid or invalid against beyond its own character rules. `translations` reports this Version's own translation status per configured language (see [Posts and publishing](/_docs/posts)'s "Translation" section) — export-only, informational: importing a document never restores or recreates this array, an imported Post simply starts with nothing translated yet, as if newly Published. `sourceProvenance` (formatVersion 5+) is this Version's own source-provenance envelope — see "Source provenance (social imports)," below; optional, `null`/omitted means none is recorded.

### Source provenance (social imports)

Any Version — most usefully one of the three social PostTypes' own imported Versions — may carry a `sourceProvenance` object recording where it actually came from:

```json
{
  "platform": "facebook",
  "sourceAccountId": "1234567890",
  "sourceAccountName": "Martin Højriis",
  "sourcePostId": "9876543210",
  "sourceUrl": "https://facebook.com/martin.hoejriis/posts/9876543210",
  "originalPublishedAt": "2020-06-14T18:30:00.000Z",
  "originalTimezone": "Europe/Copenhagen",
  "originalVisibility": "public",
  "importedAt": "2026-08-15T00:00:00.000Z",
  "converterName": "social2kartotek",
  "converterVersion": "1.0.0",
  "sourceChecksum": "b7e23ec29af22b0b4e41da31e868d57226121c84"
}
```

Every field is optional — missing optional source metadata never blocks import, whether the whole object is absent or just some fields within it. `platform` must be `"facebook"`, `"instagram"`, or `"linkedin"` when present; an unrecognized value is dropped with a warning, same leniency as every other field. `sourceAccountId` + `sourcePostId` (alongside `platform`) is the preferred stable identity for duplicate detection — see "Inventory endpoint," below; `sourceUrl`/`sourceChecksum` are fallbacks a converter may use when a stable id isn't available. On import, a Version submitting no `sourceProvenance` key at all carries forward the previous Version's own stored value unchanged (an ordinary content edit never silently wipes source facts) — but every Version this doc's own export/import machinery builds from a document always sets the key explicitly (to `null` when the source document has none), so a full document reconstructs exactly what it declared, never accidentally inheriting a value from elsewhere in the same import.

### Examples

**Minimal Image post**, one Version, one ContentElement:

```json
{
  "format": "kartotek-post", "formatVersion": 1,
  "post": { "uid": "a-photo" },
  "contentElements": [
    { "key": "image:1", "kind": "image", "role": "cover", "rank": 1, "reference": null, "createdAt": "2026-07-19T00:00:00.000Z",
      "data": { "mode": "file", "contentUrl": "https://example.com/photo.jpg", "embeddedData": null, "caption": "A photo", "metadata": null } }
  ],
  "versions": [
    { "ordinal": 1, "createdAt": "2026-07-19T00:00:00.000Z", "postType": "image", "title": "A Photo", "hidden": false, "hiddenReason": null, "content": ["image:1"] }
  ]
}
```

**Article with Text and Image**, then a second Version that edits the text and carries the same image forward (structural sharing — `image:1` is referenced by both Versions, defined only once):

```json
{
  "format": "kartotek-post", "formatVersion": 1,
  "post": { "uid": "an-article", "categoryUid": "main" },
  "contentElements": [
    { "key": "text:1", "kind": "text", "role": "body", "rank": 1, "reference": null, "createdAt": "2026-07-19T00:00:00.000Z", "data": { "bodyMarkdown": "First draft." } },
    { "key": "image:1", "kind": "image", "role": "cover", "rank": 1, "reference": null, "createdAt": "2026-07-19T00:00:00.000Z", "data": { "mode": "file", "contentUrl": "https://example.com/cover.jpg", "embeddedData": null, "caption": null, "metadata": null } },
    { "key": "text:2", "kind": "text", "role": "body", "rank": 1, "reference": null, "createdAt": "2026-07-19T01:00:00.000Z", "data": { "bodyMarkdown": "Revised text." } }
  ],
  "versions": [
    { "ordinal": 1, "createdAt": "2026-07-19T00:00:00.000Z", "postType": "article", "title": "An Article v1", "hidden": false, "hiddenReason": null, "content": ["text:1", "image:1"] },
    { "ordinal": 2, "createdAt": "2026-07-19T01:00:00.000Z", "postType": "article", "title": "An Article v2", "hidden": false, "hiddenReason": null, "content": ["text:2", "image:1"] }
  ]
}
```

### UID and Category rules

The requested `post.uid` is validated and checked for availability **before any media is fetched** — against current Post/Category uids, every historical alias (see [Domains, UIDs and addresses](/_docs/addressing)), and any other import currently in progress for that same uid. A conflict is a hard error (`409`); no job is created and nothing is fetched. The admin upload flow lets the uid be edited before starting; an external API caller must submit a different uid and retry. `post.categoryUid` is resolved against existing Categories — if it doesn't match one, the import proceeds uncategorized with a warning rather than failing outright, and never auto-creates a Category.

### Import safety

An imported Post:

- is always **new** — importing never merges into or overwrites an existing Post;
- always starts **PrePub**, regardless of the document's `post.stage` — nothing is auto-published;
- is reviewed through the normal admin editor before its Stage can change;
- has its declared `post.sourceLanguage` preserved exactly, and **never triggers translation of any kind** — import, re-import, and a later Stage change to Published all make zero requests to the external translation provider (see [Posts and publishing](/_docs/posts)'s "Translation" section: every translation is now an explicit, separate, Full-Admin action, never an automatic side effect of anything). This specific guarantee — zero provider calls across a representative multi-Post import, re-import, and post-import publish — is proven by an automated test running a real mock provider and asserting it never receives a single request, not just a code-review claim; it's part of this repository's standard automated test run, so a regression here would fail every build.

### External import API

For automated publishing, the same import machinery is reachable over HTTP with a scoped bearer token instead of an admin session. Create one from the admin **Imports** section's **Import Tokens** area — the secret is shown once, at creation, and never again; only its hash is stored. A token expires 12 months after it is minted (the list shows the date; a caller supplying its own `expiresAt` to the mint endpoint may set it sooner, never later, and a malformed value is refused) — plan to mint a replacement before then, since there is no scheduled rotation and the token's own expiry is the only bound on a compromise nobody notices. A token is scoped to `posts:import` — it can create/manage its own import jobs and query the deduplication inventory below; it cannot read or edit existing Posts' content, browse the database, or manage other tokens.

#### Endpoints

| Method & path | Purpose |
|---|---|
| `POST /_api/import/posts` | Submit a document. Body: `{ "document": {...}, "mediaMode": "automatic", "autoFinalize": true }`. |
| `GET /_api/import/posts/:id` | Fetch the current report for a job. |
| `PATCH /_api/import/posts/:id/media/:key` | Replace a failed media item's source URL: `{ "url": "…" }`. |
| `PUT /_api/import/posts/:id/media/:key/file` | Upload replacement bytes directly for a failed media item. |
| `POST /_api/import/posts/:id/retry` | Retry every currently-failed media item. |
| `POST /_api/import/posts/:id/finalize` | Finalize: `{ "allowMissingMedia": false }`. |
| `DELETE /_api/import/posts/:id` | Abandon the job and delete any staged media files. |
| `GET /_api/import/inventory?platform=<facebook\|instagram\|linkedin>` | List already-imported Posts for one platform — see "Inventory endpoint," below. |

All accept either the admin session cookie or `Authorization: Bearer <token>`. A token may only ever access an import job it created itself — not another token's, and not one created through the admin UI; the inventory endpoint has no such per-job scoping (there's no "job" to own), any valid token or admin session may query it.

### Inventory endpoint (external converter deduplication)

Before submitting a document, a converter can check what's already been imported for a given platform, to avoid creating a duplicate Post:

```bash
curl -H "Authorization: Bearer kit_xxxxxxxx…" \
  "https://hoejriis.dk/_api/import/inventory?platform=facebook&limit=50"
```

```json
{
  "items": [
    {
      "uid": "a-facebook-post",
      "postType": "facebook_post",
      "stage": "prepub",
      "platform": "facebook",
      "sourceAccountId": "1234567890",
      "sourceAccountName": "Martin Højriis",
      "sourcePostId": "9876543210",
      "sourceUrl": "https://facebook.com/martin.hoejriis/posts/9876543210",
      "originalPublishedAt": "2020-06-14T18:30:00.000Z",
      "sourceChecksum": "b7e23ec29af22b0b4e41da31e868d57226121c84",
      "latestVersionId": 42,
      "updatedAt": "2026-08-15T00:00:00.000Z"
    }
  ],
  "nextCursor": null
}
```

- `platform` (required) is one of `facebook`, `instagram`, `linkedin` — it selects that platform's own social PostType (`facebook_post`/`instagram_post`/`linkedin_post`), so every Post of that Type is included even when it has no recorded `sourceProvenance` at all (missing optional identity is represented as `null` per field, never omitted or fabricated).
- Results include Posts in **every** Stage — Published, PrePub, Hidden, and Private — with `stage` reported on each record, never filtered out. This is deliberately different from every other public/admin listing in this app, which excludes non-Published content by default; the inventory's whole purpose is duplicate avoidance before submission, which requires seeing everything already imported regardless of visibility.
- `limit` (optional, default 50, max 200) and `cursor` (optional, opaque — pass back the previous response's own `nextCursor`) paginate deterministically by Post id ascending. `nextCursor` is `null` once the last page has been returned. Repeated calls against unchanged data return identical pages; a Post created or edited between calls can't cause an already-issued page to skip or duplicate a record.
- The endpoint reports inventory only — it never decides whether two records are duplicates. The preferred identity match is `platform` + `sourceAccountId` + `sourcePostId`; `sourceUrl` or `sourceChecksum` may be used conservatively when a stable id isn't available. Multiple Posts sharing the same identity are all returned (never deduplicated or hidden) so the converter can detect and handle the ambiguity itself.
- Never exposes full Post content, media, or fields belonging to any other Post Type — only the identity/diagnostic fields shown above.

#### Example

```bash
curl -X POST https://hoejriis.dk/_api/import/posts \
  -H "Authorization: Bearer kit_xxxxxxxx…" \
  -H "Content-Type: application/json" \
  -d '{"document": {"format":"kartotek-post","formatVersion":1,"post":{"uid":"external-post"}, "contentElements":[...], "versions":[...]}, "mediaMode": "automatic", "autoFinalize": true}'
```

#### Responses

- **`201 Created`** — validated, media acquired, and finalized immediately (only when `autoFinalize` was requested and nothing failed). Body includes `resultPostId` and the full report.
- **`202 Accepted`** — the job is valid but staged for review (a media item failed, or `autoFinalize` wasn't requested). Body is the report; use the endpoints above to correct and finalize.
- **`400 Bad Request`** — the document itself is structurally invalid (bad format/version, invalid Role, a Version referencing an unknown ContentElement key, and so on). No job is created.
- **`409 Conflict`** — the requested uid is already in use, or another import for it is already in progress.
- **`401`** — no valid session or token. **`429`** — a token exceeded its rate limit.

### Security

Every media acquisition — whether from the admin UI or the external API — goes through the same protections normal media ingestion already has (see [Media and the Media Archive](/_docs/media)): per-kind size limits, an allowed content-type list, no redirect-following, and SSRF-safe URL fetching (private/loopback/link-local addresses rejected, including cloud metadata addresses). Token secrets are never logged or returned again after creation; a revoked or expired token is rejected the same generic way as a wrong one.

### Format versioning

`formatVersion` exists so this format can evolve without breaking older exports. This deployment currently understands versions `1` through `5` — a document declaring any other version is rejected outright rather than guessed at. Version 2 added `metadataEnvelope` (Text/Image/Link/Geospatial, extended to Video/Audio later — still additive, no formatVersion bump needed for that extension) and `format` (Text) — both purely additive: a version-1 document simply doesn't have them, which imports exactly as before (Text defaults to `format: "markdown"`, every envelope comes in `null`). Version 3 changed Contact's own shape: a single `contactKind`/`value` pair became an `entries[]` array — an older document's single-pair shape still imports, becoming a one-entry array. Version 4 changed Timeline's own shape: a single `dateStart`/`dateEnd`/`label`/`description` element became a container (`title`/`precision`/`events[]`) — an older document's single-date shape still imports, becoming a one-Event container with Precision `"full"` (see the `timeline` row above). Version 5 added three new `postType` values (`facebook_post`/`instagram_post`/`linkedin_post`) and each Version's own optional `sourceProvenance` object — both purely additive: a pre-5 document simply has neither, which imports exactly as before (no social-typed Version to begin with, `sourceProvenance` comes back `null`). The two full examples above are deliberately still `formatVersion: 1`, to show that an old-style document keeps working unchanged. A future version bump would document its own migration path here, not silently reinterpret an old document's fields.

## Media import (Archive Media)

Archive Media can also be delivered from outside Kartotek, through its own versioned JSON format, `kartotek-media`. An external producer does all the preparation (finding the originals, reading their metadata, classifying them, making a ready-to-serve thumbnail) and pushes finished records here in small batches. Kartotek checks, stores and lists what it is given. It never fetches an original, reads a file's EXIF or XMP data, classifies an image or resizes one on the producer's behalf, and it never reaches out to the producer or to where the originals are stored.

An item this import has accepted is **import-managed**: from then on its metadata and thumbnail come only from the producer. Kartotek has no scanning or backfill job of its own that would touch an Archive Media item — this import is the only way one arrives or changes. See [Media and the Media Archive](/_docs/media)'s "Archive Media" section for how imported values sit alongside the ones you set yourself.

### Credential

Mint a token from the admin **Imports** section's **Import Tokens** area and choose the **Media import** scope (`media:import`). The secret is shown once. A media token can submit media batches, read its own batch reports and read the media inventory, and nothing else. A `posts:import` token or an agent token is refused here exactly like a wrong one. An admin session is also accepted, which is what the Imports section's **Media imports** panel uses (paste a batch, optionally tick **Dry run**, submit, read the report).

### Endpoints

| Method & path | Purpose |
|---|---|
| `POST /_api/import/media` | Submit a batch. Body: `{ "batch": {...}, "dryRun": false }`. |
| `GET /_api/import/media/:id` | A batch's recorded report. A token reads only the batches it submitted. |
| `GET /_api/import/media/inventory?limit=&cursor=` | What Kartotek holds, for reconciliation (below). |

Responses: **`200`**, the batch was processed; every item's own outcome is in the report, including items that failed. **`400`**, the batch itself is invalid (wrong `format`/`formatVersion`, no items, more than 50 items); nothing was written. **`401`**, no valid session or media token. **`413`**, the request body is over 64 MB. **`429`**, the token's rate limit (30 requests a minute). **`507`**, the server lacks the free disk space to store the batch's assets; nothing was written.

### Batch shape

```json
{
  "batch": {
    "format": "kartotek-media",
    "formatVersion": 1,
    "producer": { "name": "my-producer", "version": "1.4.0" },
    "items": [ { "…one item, below…": true } ]
  },
  "dryRun": false
}
```

`producer` is informational and is recorded with the batch. `dryRun: true` validates every item and reports the outcome each would get, without writing anything.

### Item shape

```json
{
  "source": {
    "accountId": "dbid:AAB…",
    "fileId": "id:a4ayc_80_OEAAAAAAAAAXw",
    "root": "/Billeder/Temaer",
    "path": "/Billeder/Temaer/Natur/IMG_0001.jpg",
    "rev": "015f…",
    "contentHash": "e3b0…",
    "modifiedTime": "2024-06-01T18:31:02Z",
    "byteSize": 4812734,
    "checksumSha256": "9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08",
    "state": "available"
  },
  "technical": { "mimeType": "image/jpeg", "width": 4032, "height": 3024, "orientation": 1, "lens": "…", "focalLength": "4.2 mm" },
  "revision": 3,
  "mode": "full",
  "fields": {
    "title": { "value": "Evening run along the harbour", "source": "owner", "writtenAt": "2026-09-01T10:00:00Z" },
    "captureTime": { "value": { "raw": "2024-06-01T18:30:00", "offset": "+02:00" }, "source": "exif" },
    "location": { "value": { "latitude": 56.15, "longitude": 10.21, "city": "Aarhus", "countryName": "Denmark", "countryCode": "DK", "provinceState": null, "sublocation": null }, "source": "exif" },
    "environment": { "value": "outdoor", "source": "classifier", "confidence": 0.93, "model": "scene-v2" },
    "tags": { "value": ["harbour", "evening"], "source": "owner" },
    "quality": { "value": "high", "source": "owner" }
  },
  "assets": {
    "thumbnail": { "contentType": "image/jpeg", "byteLength": 48211, "sha256": "…64 hex…", "data": "/9j/4AAQSkZJRg…" }
  }
}
```

**Identity.** `source.accountId` plus `source.fileId` is the item's identity: the same pair the catalogue already uses for every Archive Media item. A path, a filename or a content hash is never identity. An item whose identity already exists in the catalogue (including one Kartotek's own scan found) is updated in place and keeps its id, Categories, Tags and your own edits. Any other identity creates a new item. The same identity twice in one batch is refused for both copies.

**`source`.** `root` must be the path of an existing, non-removed Archive Root under the same account, and `path` must lie under it; an item naming any other root fails. A changed `path` for a known identity is a move: same item, new path. `state` is `available` or `missing`. `missing` is the explicit way to say the original is gone; the record stays, marked missing. An item left out of a batch is never touched, and leaving it out never means it was deleted. `rev`, `contentHash`, `modifiedTime`, `byteSize` and `checksumSha256` describe the original and are stored when given, left as they are when not.

**`technical`.** Optional descriptive facts about the original, stored when given.

**`revision`.** A positive integer the producer raises every time it changes anything about the item. Kartotek remembers the last revision it accepted for each item, together with a hash of that delivery's content, and decides each item's outcome from them (below).

**`mode`.** `full` (the default) is a complete statement of the item's fields: a field the producer delivered before and leaves out now is removed from the item's imported values. `patch` changes only the fields it lists, and a field given as `null` removes that one field. In full mode, `null` is invalid; leave the field out instead.

**`fields`.** Up to twelve fields, each an object with `value`, `source`, and optionally `confidence` (0 to 1), `model`, `method` and `writtenAt`:

| Field | `value` | Allowed `source` |
|---|---|---|
| `captureTime` | `{ "raw": "YYYY-MM-DDTHH:MM:SS", "offset": "+02:00" or null, "precision": null }`: the camera's wall clock, never converted to UTC. A reduced-precision time sets `precision` to `"year"` or `"month"`, with `raw` truncated to `YYYY` or `YYYY-MM` (never padded) and no offset; it is kept as the imported value but never becomes the item's capture time, whatever its source | `exif`, `draft`, `owner` |
| `exposure` | `{ "exposureTime", "fNumber" }`, strings or null | `exif` |
| `device` | `{ "make", "model" }`, strings or null | `exif`, `owner` |
| `location` | `{ "latitude", "longitude", "city", "provinceState", "countryName", "countryCode", "sublocation" }`, decimal degrees; latitude and longitude both given or both null | `exif`, `inferred`, `draft`, `owner` |
| `environment` | `indoor`, `outdoor`, `mixed`, `uncertain` | `classifier`, `owner` |
| `people` | `none_detected`, `people_not_identifiable`, `potentially_identifiable`, `uncertain` | `classifier`, `owner` |
| `privacy` | `private`, `public` | `draft`, `owner` |
| `quality` | `low`, `medium`, `high` | `draft`, `owner` |
| `title`, `description`, `notes` | a non-empty string | title/description: `draft`, `owner`; notes: `owner` |
| `tags` | an array of Tag spellings | `draft`, `owner` |

The source decides whether a value takes effect. `draft` is a proposal: it is stored and shown in the item's detail but never becomes the item's value. Title, description, notes, tags and quality take effect only from `owner`. Capture time and device take effect from `exif` or `owner`, location from `exif`, `inferred` or `owner`, and environment/people from `classifier` or `owner`. **Privacy never takes effect from an import**, whatever its source. The producer's privacy value is kept for display, but making an item public is always a choice made in Kartotek, so no import can publish anything.

**`assets.thumbnail`.** The ready-to-serve thumbnail every Archive Media view uses: JPEG only, at most 8 MB, the bytes base64-encoded in `data` (no `data:` prefix), with the exact `byteLength` and lowercase hex `sha256` of the decoded bytes. All three are checked and the item fails on any mismatch. A replacement thumbnail takes over only once the item's new revision is accepted, and the previous one keeps serving until then; if storing it fails, the item fails and the old thumbnail stays. An item delivered without a thumbnail keeps the one it has. A brand-new item without one shows as awaiting a thumbnail, and Kartotek never generates one. Omit `assets` entirely when nothing changed; there is no URL form, since Kartotek never fetches anything.

### Outcomes

Each item in the report has an `outcome`:

- **`created`**: a new catalogue item.
- **`updated`**: a higher revision than the last accepted one, or the first import of an item Kartotek already had (reported with `"adopted": true`).
- **`unchanged`**: the same revision with the same content as the last accepted delivery. Nothing is written, so replaying a batch is safe.
- **`conflict`**: refused, nothing written. `reason` is `stale_revision` (older than the accepted revision), `revision_reused` (the accepted revision number with different content) or `duplicate_in_batch`.
- **`failed`**: refused, nothing written. `reason` is `invalid` (see `errors`, each with a `path` into the batch), `unknown_root` or `asset_write_failed`. Fix it and resend; the same revision is accepted then.

Items are independent. One failing never affects another, and each is written all-or-nothing. Two deliveries of the same item racing each other are serialized: one is accepted and the other gets the outcome it would have had afterwards. The report also carries each item's `archiveItemId` and whether its thumbnail was `stored`, `unchanged` or `absent`, and the batch's `counts` per outcome.

### Imported values and your own edits

An imported value never overwrites a value you set in Kartotek. If the producer delivers title A, you change it to B, and a later delivery brings C, the item still shows B, and its detail panel shows C as the imported value. Choosing to follow the imported value again switches it to C and to whatever later deliveries bring. Categories are not part of this format: an imported item gets a Category only when you assign one yourself in Kartotek (see [Organizing and presenting content](/_docs/categories)'s "Categories and Tags on media"), and Categories and Tags you add by hand are never touched by an import.

### Inventory

`GET /_api/import/media/inventory` lists every Archive Media item, imported or not, in ascending id order: `limit` defaults to 100 (max 200), and `nextCursor` is passed back as `cursor` until it is `null`. Pages over unchanged data are identical, and an item added between calls never makes a page skip or repeat one. Each entry has only reconciliation fields: `id`, `accountId`, `fileId`, `root`, `rootStatus`, `path`, `sourceState`, `checksumSha256`, `revision` and `payloadSha256` (the last accepted delivery; both `null` for an item no import has touched yet), `acceptedAt`, and `thumbnail` (`state`, `sha256`). No titles, descriptions, notes, Tags or classification. A producer reconciles by comparing identity and revision, then sends only what differs.

### Limits

50 items per batch, 64 MB per request body, 8 MB per thumbnail, 30 requests a minute per token. Base64 adds a third to an asset's size, so batches of large thumbnails should be smaller. Every batch checks free disk space before it writes anything.

### JSON Schema (format version 1)

```json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "kartotek-media batch, format version 1",
  "type": "object",
  "required": ["batch"],
  "properties": {
    "dryRun": { "type": "boolean" },
    "batch": {
      "type": "object",
      "required": ["format", "formatVersion", "items"],
      "properties": {
        "format": { "const": "kartotek-media" },
        "formatVersion": { "const": 1 },
        "producer": { "type": "object", "properties": { "name": { "type": "string" }, "version": { "type": "string" } } },
        "items": { "type": "array", "minItems": 1, "maxItems": 50, "items": { "$ref": "#/$defs/item" } }
      }
    }
  },
  "$defs": {
    "item": {
      "type": "object",
      "required": ["source", "revision"],
      "properties": {
        "source": {
          "type": "object",
          "required": ["accountId", "fileId", "root", "path"],
          "properties": {
            "accountId": { "type": "string", "minLength": 1 },
            "fileId": { "type": "string", "minLength": 1 },
            "root": { "type": "string" },
            "path": { "type": "string", "minLength": 1 },
            "rev": { "type": ["string", "null"] },
            "contentHash": { "type": ["string", "null"] },
            "modifiedTime": { "type": ["string", "null"], "format": "date-time" },
            "byteSize": { "type": ["integer", "null"], "minimum": 0 },
            "checksumSha256": { "type": ["string", "null"], "pattern": "^[0-9a-fA-F]{64}$" },
            "state": { "enum": ["available", "missing"] }
          }
        },
        "technical": {
          "type": "object",
          "properties": {
            "mimeType": { "type": ["string", "null"] },
            "width": { "type": ["integer", "null"], "minimum": 1 },
            "height": { "type": ["integer", "null"], "minimum": 1 },
            "orientation": { "type": ["integer", "null"], "minimum": 1 },
            "lens": { "type": ["string", "null"] },
            "focalLength": { "type": ["string", "null"] }
          }
        },
        "revision": { "type": "integer", "minimum": 1 },
        "mode": { "enum": ["full", "patch"] },
        "fields": {
          "type": "object",
          "propertyNames": { "enum": ["captureTime", "exposure", "device", "location", "environment", "people", "privacy", "quality", "title", "description", "notes", "tags"] },
          "additionalProperties": { "oneOf": [{ "type": "null" }, { "$ref": "#/$defs/field" }] }
        },
        "assets": {
          "type": "object",
          "additionalProperties": false,
          "properties": {
            "thumbnail": {
              "type": ["object", "null"],
              "required": ["contentType", "byteLength", "sha256", "data"],
              "properties": {
                "contentType": { "const": "image/jpeg" },
                "byteLength": { "type": "integer", "minimum": 1, "maximum": 8388608 },
                "sha256": { "type": "string", "pattern": "^[0-9a-f]{64}$" },
                "data": { "type": "string", "contentEncoding": "base64" }
              }
            }
          }
        }
      }
    },
    "field": {
      "type": "object",
      "required": ["value", "source"],
      "properties": {
        "value": {},
        "source": { "enum": ["exif", "classifier", "inferred", "draft", "owner"] },
        "confidence": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
        "model": { "type": ["string", "null"] },
        "method": { "type": ["string", "null"] },
        "writtenAt": { "type": ["string", "null"], "format": "date-time" }
      }
    }
  }
}
```

The per-field `value` shapes and allowed sources in the table above are checked as well; the schema leaves them to that table.

### Minimal example

A new item with just a title and a thumbnail:

```bash
curl -X POST https://hoejriis.dk/_api/import/media \
  -H "Authorization: Bearer kit_xxxxxxxx…" \
  -H "Content-Type: application/json" \
  -d '{"batch":{"format":"kartotek-media","formatVersion":1,"items":[{"source":{"accountId":"dbid:AAB…","fileId":"id:a4ayc…","root":"/Billeder/Temaer","path":"/Billeder/Temaer/Natur/IMG_0001.jpg"},"revision":1,"fields":{"title":{"value":"Evening run","source":"owner"}},"assets":{"thumbnail":{"contentType":"image/jpeg","byteLength":48211,"sha256":"…","data":"/9j/…"}}}]}}'
```

The full example is the item shape above. Changing only the title later is a patch:

```json
{ "source": { "accountId": "dbid:AAB…", "fileId": "id:a4ayc…", "root": "/Billeder/Temaer", "path": "/Billeder/Temaer/Natur/IMG_0001.jpg" },
  "revision": 4, "mode": "patch", "fields": { "title": { "value": "Evening run, June", "source": "owner" } } }
```

## See also

[Posts and publishing](/_docs/posts) for the Post/Version/ContentElement model this format serializes, [Media and the Media Archive](/_docs/media) for how Image/Video/Audio ingestion and its security limits work, [Domains, UIDs and addresses](/_docs/addressing) for uid/alias rules, and [Administrator quick start](/_docs/admin-quickstart) for the admin editor generally.
