Developer docs

Quickstart

Fully OpenAI-compatible — change one base_url and you are in. Existing code keeps its structure, and auth, rate limiting, streaming and usage metering all follow the standard semantics.

Before you start

Three steps to your first call. Nothing to request, no instance type to choose, and no need to care which silicon runs underneath.

  1. Create an account. Sign up with an email and password on the registration page; a default key is generated for you.
  2. Create an API key in the console. Open the API Key page in the console and add one. Keys can be organised per environment — production, staging, development — and disabled at any time.
  3. Note the base_url. Point your SDK at the address below; every other parameter stays exactly as it is with the official SDK.
base_url
https://api.yrinvestment.net/v1
One address covers every model. The same base_url and the same key reach chat, multimodal, embedding, reranking and speech models alike; switching model is a change to the model field and nothing else. The live catalog is at the Model Catalog.

Authentication

Every endpoint uses HTTP Bearer auth: put the API key in the Authorization header. This matches OpenAI exactly, so the official SDKs attach it for you once api_key is set.

header
Authorization: Bearer sk-yr-3f8a1c2d7b4e5a90c1d2e3f4a5b6c7d8e9f0a1b2

Anatomy of a key

A key has two parts. The prefix is sk-yr- plus eight hex characters, which identifies the key; a random secret follows. The console and the call logs show only the prefix, so you can tell which key is in use.

The full key is shown once, at creation. Once the dialog closes it cannot be retrieved, so store it in a secret manager or an environment variable straight away. If it is lost, the only path is to disable the old key and create a new one.

Never put a key in front-end code or commit it to a repository. Keep it in an environment variable such as YR_API_KEY, and split it into several keys by environment — rate limit, quota and the set of callable models are configured per key, so a problem key can be disabled on its own without touching anything else.

Chat completions POST /v1/chat/completions

The endpoint you will use most. Semantics are identical to OpenAI Chat Completions. The platform dispatches the request, per the compatibility matrix and the routing policy, to silicon that can run the model and has capacity right now.

Request parameters

ParameterTypeRequiredDescription
modelstringRequired Model name, e.g. DeepSeek-V4.1-Flash or Qwen3-32B. See the Model Catalog or /v1/models.
messagesarrayRequired Array of chat messages. Each carries a role (system / user / assistant) and content.
streambooleanOptional Whether to stream, default false. Set true to receive Server-Sent Events, described in the next section.
temperaturenumberOptional Sampling temperature, 0 to 2, default 1. Lower is more deterministic — better for extraction and classification.
max_tokensintegerOptional Maximum tokens to generate. If omitted, it follows the model context length and the space left in it.

Example call

request.sh
curl https://api.yrinvestment.net/v1/chat/completions \
  -H "Authorization: Bearer $YR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "DeepSeek-V4.1-Flash",
    "messages": [
      {"role": "system", "content": "You are a precise technical assistant."},
      {"role": "user", "content": "In one sentence, what problem does heterogeneous compute scheduling solve?"}
    ],
    "temperature": 0.6,
    "max_tokens": 512,
    "stream": false
  }'
quickstart.py
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["YR_API_KEY"],
    base_url="https://api.yrinvestment.net/v1",  # change only this line
)

resp = client.chat.completions.create(
    model="DeepSeek-V4.1-Flash",
    messages=[
        {"role": "system", "content": "You are a precise technical assistant."},
        {"role": "user", "content": "In one sentence, what problem does heterogeneous compute scheduling solve?"},
    ],
    temperature=0.6,
    max_tokens=512,
)

print(resp.choices[0].message.content)
print(resp.usage.total_tokens)
print(resp.model_extra["yr_routing"])  # which silicon and which compute centre served this call
quickstart.mjs
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.YR_API_KEY,
  baseURL: "https://api.yrinvestment.net/v1",
});

const resp = await client.chat.completions.create({
  model: "Qwen3-32B",
  messages: [
    { role: "system", content: "You are a precise technical assistant." },
    { role: "user", content: "In one sentence, what problem does heterogeneous compute scheduling solve?" },
  ],
  temperature: 0.6,
  max_tokens: 512,
});

console.log(resp.choices[0].message.content);
console.log(resp.usage.total_tokens);
console.log(resp.yr_routing);
Responses carry one extra field, yr_routing. It is a platform extension on top of the standard response; the standard SDKs drop it into their unknown-field container (model_extra in Python, a plain property in Node) and nothing about the official semantics changes.

Streaming

Set stream to true and the endpoint switches to Server-Sent Events: the response header is text/event-stream, the body is a series of data: lines, each one a JSON chunk, and the text delta sits in choices[0].delta.content. The final chunk carries finish_reason and usage, followed by data: [DONE].

response (text/event-stream)
data: {"id":"chatcmpl-8f2a1c7d4b9e0a3f5c6d","object":"chat.completion.chunk","created":1789243100,"model":"DeepSeek-V4.1-Flash","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}],"yr_routing":{"chip":"h200","chip_name":"H200","site":"HK-TKO","ttft_ms":180}}

data: {"id":"chatcmpl-8f2a1c7d4b9e0a3f5c6d","object":"chat.completion.chunk","created":1789243100,"model":"DeepSeek-V4.1-Flash","choices":[{"index":0,"delta":{"content":"Heterogene"},"finish_reason":null}]}

data: {"id":"chatcmpl-8f2a1c7d4b9e0a3f5c6d","object":"chat.completion.chunk","created":1789243100,"model":"DeepSeek-V4.1-Flash","choices":[{"index":0,"delta":{"content":"ous sched"},"finish_reason":null}]}

data: {"id":"chatcmpl-8f2a1c7d4b9e0a3f5c6d","object":"chat.completion.chunk","created":1789243100,"model":"DeepSeek-V4.1-Flash","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":42,"completion_tokens":168,"total_tokens":210}}

data: [DONE]

On the first chunk: yr_routing

A platform-specific extension field, present on the first streamed chunk (and at the top level of a non-streamed response). It tells you where the call actually landed — the silicon underneath is heterogeneous, but you never have to guess which card you got.

FieldExampleMeaning
chiph200Identifier of the silicon family that served this inference.
chip_nameH200Display name of that silicon family — H200, H100, Ascend A2, Hygon DCU.
siteHK-TKOCompute centre code, useful for proximity routing and for checking data residency.
ttft_ms180Time to first token for this call, in milliseconds.

If the workload has hard requirements on jurisdiction or silicon family — domestic cards only, data must not leave the country — set a routing policy on the relevant key in the console, then use this field to verify it client-side.

stream.py
stream = client.chat.completions.create(
    model="DeepSeek-V4.1-Flash",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

for chunk in stream:
    routing = getattr(chunk, "yr_routing", None)
    if routing:
        print("served by", routing["chip_name"], "@", routing["site"])
    print(chunk.choices[0].delta.content or "", end="")

Model list GET /v1/models

Returns every model currently live, in the same shape as the OpenAI models endpoint: object is list, data is the model array, id is the model name you drop straight into the model field, and owned_by is the vendor. Use it to pull the selectable models at application start instead of hard-coding them.

request.sh
curl https://api.yrinvestment.net/v1/models \
  -H "Authorization: Bearer $YR_API_KEY"
response.json
{
  "object": "list",
  "data": [
    { "id": "DeepSeek-V4.1-Flash", "object": "model", "owned_by": "DeepSeek", "created": 1756000000 },
    { "id": "Qwen3-32B",  "object": "model", "owned_by": "Qwen", "created": 1756000000 }
  ]
}

Which silicon families each model runs on, and what the evidence is, is set out in the compatibility matrix.

Rate limits & quota

Each API key carries three independent limits. They do not substitute for one another, and hitting any one of them rejects the request. The current level of all three is live in the console.

LimitCounted asDefaultOn breach
RPMRequests per minute, counted per key on a sliding window 3,000 Returns 429 — retry with exponential backoff
TPMTokens per minute, input and output combined 8,000,000 Returns 429 — split the batch or lower concurrency
Monthly token quotaTotal tokens the key has consumed within the calendar month Per key Returns 401 — raise the quota in the console or move to another key
Split keys by workload. RPM and TPM are both counted per key, so putting online traffic and offline batch work on separate keys means a batch job saturating its own limit cannot crowd out live requests. The set of callable models can be narrowed per key as well.

Error codes

Errors return JSON in a fixed shape: {"error": {"message": "...", "code": 401}}. message is written to be shown to an operator as-is, and code matches the HTTP status.

StatusMeaningTypical cause and fix
400Malformed request Missing model or messages, a wrong field type, or unparseable JSON. Check the body and resend.
401Auth failed or quota spent Missing Authorization header, invalid key, disabled key, or the monthly token quota for that key is spent.
403Forbidden The account role cannot access this view, or the key is not allowed to call this model.
404Endpoint or resource not found Wrong path, or the requested model is not live. Check the name against /v1/models.
429Rate limited RPM or TPM exceeded. Retry with exponential backoff, lower concurrency, or split the batch.
500Internal server error A platform-side fault. Send us the request id from the response and we can pin it to the exact scheduling record.