Developer Documentation

Augure API

OpenAI-compatible chat completions API with Canadian data residency. Inference in Canada — never US providers.

Base URL: https://api.augureai.ca
🍁

Data Routing & Residency

All API requests enter through our gateway in Beauharnois, Quebec. Inference runs in Canada under zero-data-retention agreements — never routed to US providers. Image reading and voice synthesis run on GPU servers Augure manages in Calgary, Alberta. Prompts are encrypted in transit (TLS 1.2+), never logged by Augure, and never used for model training.

Canadian gatewaySovereign inferenceNever US providersNo prompt logging

Authentication

All API endpoints require a Bearer token. Include your API key in the Authorization header of every request.

Example
curl https://api.augureai.ca/v1/models \
  -H "Authorization: Bearer YOUR_API_KEY"

Getting a key: API keys are issued through our gated application process. Get API access to get started.

Models

Five models are available, optimized for different workloads, plus an auto alias that routes each request to the best available model.

ossington-5

Newest flagship reasoning model. Fastest on the menu, strong at code.

Agentic coding, complex reasoning, structured output

Always on

rosedale-1

Premium reasoning tier. Extended thinking for hard, multi-step problems.

Advanced reasoning, agentic workflows, complex analysis

Always on

ossington-4

Large multimodal model, high capability

Complex reasoning, legal analysis, document review

Always on

ossington-4-1

Premium reasoning on Canadian infrastructure. Tools, strong bilingual (EN/FR) work.

Complex reasoning, legal analysis, document review

Always on

tofino-3

Light tier. Fast, thinking off, served in Canada with no failover outside Canada.

Chat, summaries, quick tasks

Always on

OpenAI compatibility: The aliases gpt-4, gpt-4o, gpt-4o-mini, and gpt-3.5-turbo are supported for drop-in compatibility with OpenAI client libraries. They map to ossington-4 and tofino-3 respectively.

Endpoints

POST/v1/chat/completions

Create a chat completion. Accepts the same request format as the OpenAI chat completions endpoint.

Parameters

FieldTypeRequiredDescription
modelstringYesModel ID (see Models above)
messagesarrayYesArray of message objects (text, images or PDF files — see below)
streambooleanNoStream response via SSE. Default: false
stream_optionsobjectNo{"include_usage": true} adds a final usage chunk before [DONE]
temperaturenumberNoSampling temperature (0.0–2.0)
max_tokensnumberNoMax tokens to generate (up to 32,768)
top_pnumberNoNucleus sampling threshold
stopstring | arrayNoStop sequence(s)

Each message in the messages array has a role ("system", "user", or "assistant") and a content string.

Images and documents

content can also be an array of parts. Alongside text parts, a user message may carry image_url parts (PNG, JPEG, WebP or GIF as a base64 data URL) and file parts (a PDF as a base64 data URL). PDFs are read on the gateway: the text layer is extracted directly, and scanned documents are OCR’d automatically. Send the PDF itself rather than rendering it to an image; you keep the exact text, and it works with every model.

PDF attachment
curl -X POST https://api.augureai.ca/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ossington-5",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Extract the invoice number, date, vendor and total as JSON."},
        {"type": "file", "file": {
          "filename": "invoice-2041.pdf",
          "file_data": "data:application/pdf;base64,JVBERi0xLjQK..."
        }}
      ]
    }]
  }'
Python
import base64
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.augureai.ca/v1")
pdf = base64.b64encode(open("invoice-2041.pdf", "rb").read()).decode()

response = client.chat.completions.create(
    model="ossington-5",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Extract the invoice number, date, vendor and total as JSON."},
            {"type": "file", "file": {
                "filename": "invoice-2041.pdf",
                "file_data": f"data:application/pdf;base64,{pdf}",
            }},
        ],
    }],
)
print(response.choices[0].message.content)

Each PDF may be up to 25 MB and 200 pages, within the 35 MB request body. A file part also accepts spreadsheets (.xlsx or .csv data URLs, first 100 rows of each sheet). Images are read by our vision engine and handed to the model as text; add "augure_ocr": true to an image_url part for a verbatim OCR transcription with no description. The gateway keeps no copy of a file once the request completes.

Example Request

curl
curl -X POST https://api.augureai.ca/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ossington-5",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the Civil Code of Quebec?"}
    ]
  }'

Example Response

Response
{
  "id": "chatcmpl-a9adf17e-5ff3-4804-b01e-f7cbd30ae996",
  "object": "chat.completion",
  "created": 1771286577,
  "model": "ossington-5",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The Civil Code of Quebec (Code civil du Québec) is..."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 24,
    "completion_tokens": 150,
    "total_tokens": 174
  },
  "_augure": {
    "gateway_region": "ca-montreal-1",
    "inference_region": "augure-cloud",
    "request_id": "a9adf17e-5ff3-4804-b01e-f7cbd30ae996"
  }
}

Streaming

Set "stream": true to receive Server-Sent Events. Each event is a JSON chunk with a delta object containing incremental content. The stream ends with data: [DONE]. Add "stream_options": {"include_usage": true} to receive a final chunk carrying the usage object (with empty choices) just before it.

Streaming request
curl -N -X POST https://api.augureai.ca/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tofino-3",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'

Usage and cost

Every response carries a usage object with prompt_tokens, completion_tokens and total_tokens; multiply by the per-model rates above to price a call. Your developer dashboard shows the same numbers per model, per day and per key, with the cost in CAD at the published rates. Text extracted from an attached PDF counts toward prompt tokens like any other input.

GET/v1/models

Returns a list of all available models.

curl
curl https://api.augureai.ca/v1/models \
  -H "Authorization: Bearer YOUR_API_KEY"

Response

Response
{
  "object": "list",
  "data": [
    { "id": "auto",          "object": "model", "owned_by": "augure" },
    { "id": "rosedale-1",    "object": "model", "owned_by": "augure" },
    { "id": "ossington-4-1", "object": "model", "owned_by": "augure" },
    { "id": "ossington-5",   "object": "model", "owned_by": "augure" },
    { "id": "ossington-4",   "object": "model", "owned_by": "augure" },
    { "id": "tofino-3",      "object": "model", "owned_by": "augure" }
  ]
}

Client Libraries

Use any OpenAI-compatible SDK. Just point it to https://api.augureai.ca/v1 as the base URL.

Python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.augureai.ca/v1"
)

response = client.chat.completions.create(
    model="ossington-5",
    messages=[
        {"role": "user", "content": "Explain Quebec privacy law"}
    ]
)
print(response.choices[0].message.content)
JavaScript / TypeScript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "YOUR_API_KEY",
  baseURL: "https://api.augureai.ca/v1",
});

const response = await client.chat.completions.create({
  model: "tofino-3",
  messages: [{ role: "user", content: "Summarize PIPEDA" }],
});
console.log(response.choices[0].message.content);

Limits

Request body

35 MB max

PDF attachment

25 MB / 200 pages

Messages per request

256 max

Max output tokens

32,768

Request timeout

300 seconds

Rate limits are applied per API key; the daily spend cap applies across your account. Contact us if you need higher throughput for production workloads.

Errors

All errors return a JSON object with an error field, matching the OpenAI error format.

Error response
{
  "error": {
    "message": "Invalid API key provided",
    "type": "invalid_request_error",
    "param": null,
    "code": "invalid_api_key"
  }
}
StatusMeaning
401Missing or invalid API key
400Malformed request or missing required fields
404Unknown model or endpoint
413Request body exceeds 35 MB
429Token quota exceeded for this API key
502Upstream processing error — retry shortly

Ready to integrate?

Get your API key and start building with Augure.

Get API Access