> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sajn.se/llms.txt
> Use this file to discover all available pages before exploring further.

# Sync completed documents nightly

> Recipe: a scheduled job that downloads every newly completed document's sealed PDF and journal into your archive

In this recipe, you write a job that runs every night and copies each document completed since the last run into your archive: the sealed PDF, the signing journal, and a JSON file with the document's details. It's the safety net next to webhooks: if your webhook endpoint missed a `document.completed` delivery, the next night's run catches the document.

## Before you begin

* Python 3.9 or later, and the `requests` package: `pip install requests`.
* An API key in the `SAJN_API_KEY` environment variable. To create one, go to workspace settings in the sajn app, then **Utvecklare** (Developer) > **API-nycklar** (API keys). The job archives the documents of the key's workspace.
* A scheduler, such as cron, that can run the job once a night.

## Build the job

<Steps>
  <Step title="Write the job">
    Save the following file as `nightly_sync.py`. It reads the last run's timestamp from `last-sync.txt`, lists documents completed since then with a cursor, archives each one, and saves the new timestamp only after every document succeeded:

    ```python theme={null}
    import json
    import os
    import pathlib
    import time

    import requests

    API = "https://app.sajn.se/api/v1"
    ARCHIVE = pathlib.Path("archive")
    STATE = pathlib.Path("last-sync.txt")
    SESSION = requests.Session()
    SESSION.headers.update({
        "Authorization": f"Bearer {os.environ['SAJN_API_KEY']}",
        "Sajn-Version": "2026-10",
    })


    def get(path, params=None):
        """Sends a GET request, waiting and retrying on 429 and 5xx."""
        for attempt in range(5):
            response = SESSION.get(f"{API}{path}", params=params)
            if response.status_code == 429 or response.status_code >= 500:
                time.sleep(int(response.headers.get("Retry-After", 2 ** attempt)))
                continue
            data = response.json()
            if not response.ok:
                raise RuntimeError(f"GET {path}: {data['code']} {data['message']} ({data['requestId']})")
            return data
        raise RuntimeError(f"GET {path}: gave up after retries")


    def download(document_id, file_type, target):
        url = get(f"/documents/{document_id}/files/{file_type}")["url"]
        file = requests.get(url)
        file.raise_for_status()
        target.write_bytes(file.content)


    def archive(document):
        folder = ARCHIVE / document["id"]
        folder.mkdir(parents=True, exist_ok=True)
        details = get(f"/documents/{document['id']}")
        (folder / "document.json").write_text(json.dumps(details, indent=2, ensure_ascii=False))
        download(document["id"], "SIGNED", folder / "sealed.pdf")
        download(document["id"], "JOURNAL", folder / "journal.pdf")


    params = {
        "status": "COMPLETED",
        "orderBy": "completedAt",
        "orderDirection": "asc",
        "limit": 100,
    }
    if STATE.exists():
        params["completedAfter"] = STATE.read_text().strip()

    started_at = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
    count = 0
    while True:
        page = get("/documents", params)
        for document in page["data"]:
            archive(document)
            count += 1
        if not page["hasMore"]:
            break
        params["cursor"] = page["nextCursor"]

    STATE.write_text(started_at)
    print(f"Archived {count} documents. Next run: completedAfter={started_at}")
    ```

    The job works as follows:

    * `completedAfter` includes its boundary, and the job stores the time the run started, so a document completed during the run is archived again the next night. Writing to a fixed folder per document makes that harmless.
    * The job passes each response's `nextCursor` as `cursor` until `hasMore` is `false`. The cursor keeps the pages stable while documents change.
    * Each file `url` expires at the response's `expiresAt`, so the job downloads it right away.
    * If any document fails, the job stops before it saves the timestamp, so the next run retries the same window.
    * Each document costs three API requests, plus its share of the list pages. For your plan's limits, see [Rate limits and quotas](/api-fundamentals/rate-limits).
  </Step>

  <Step title="Run it once by hand">
    ```bash theme={null}
    python nightly_sync.py
    ```

    The first run has no `last-sync.txt`, so it archives every completed document. The output is similar to the following:

    ```text theme={null}
    Archived 214 documents. Next run: completedAfter=2026-10-02T02:00:04Z
    ```
  </Step>

  <Step title="Schedule it">
    Add a cron entry that runs the job at 02:00 every night:

    ```text theme={null}
    0 2 * * * cd /opt/sajn-archive && SAJN_API_KEY=API_KEY /usr/bin/python3 nightly_sync.py >> sync.log 2>&1
    ```

    Replace `API_KEY` with your API key, or load it from your secret store instead.
  </Step>
</Steps>

## Handle errors

* `409 INVALID_STATE` on a file: sajn hasn't produced the file yet. sajn generates the journal shortly after it seals the document. The job stops, and the next run retries the document.
* `401 UNAUTHORIZED`: the key is invalid or was revoked.
* `400 VALIDATION_FAILED`: a malformed `completedAfter` in `last-sync.txt`. The value must be an ISO 8601 date-time.

For the error shape and every code, see [Errors](/api-fundamentals/errors).

## Related guides

* [Sync documents](/guides/integrations/syncing-documents)
* [Download documents](/guides/documents/downloading-documents)
* [Replay and reconcile](/webhooks/replay-and-reconcile)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.