Fully OpenAI-compatible — change one base_url and you are in. Existing code keeps its structure, and auth, rate limiting, streaming and usage metering all follow the standard semantics.
Three steps to your first call. Nothing to request, no instance type to choose, and no need to care which silicon runs underneath.
https://api.yrinvestment.net/v1
model field and nothing else. The live catalog is at the Model Catalog.
Every endpoint uses HTTP Bearer auth: put the API key in the Authorization header. This matches OpenAI exactly, so the official SDKs attach it for you once api_key is set.
Authorization: Bearer sk-yr-3f8a1c2d7b4e5a90c1d2e3f4a5b6c7d8e9f0a1b2
A key has two parts. The prefix is sk-yr- plus eight hex characters, which identifies the key; a random secret follows.
The console and the call logs show only the prefix, so you can tell which key is in use.
The full key is shown once, at creation. Once the dialog closes it cannot be retrieved, so store it in a secret manager or an environment variable straight away. If it is lost, the only path is to disable the old key and create a new one.
YR_API_KEY,
and split it into several keys by environment — rate limit, quota and the set of callable models are configured per key, so a problem key can be disabled on its own without touching anything else.
POST /v1/chat/completionsThe endpoint you will use most. Semantics are identical to OpenAI Chat Completions. The platform dispatches the request, per the compatibility matrix and the routing policy, to silicon that can run the model and has capacity right now.
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Required | Model name, e.g. DeepSeek-V4.1-Flash or Qwen3-32B. See the Model Catalog or /v1/models. |
| messages | array | Required | Array of chat messages. Each carries a role (system / user / assistant) and content. |
| stream | boolean | Optional | Whether to stream, default false. Set true to receive Server-Sent Events, described in the next section. |
| temperature | number | Optional | Sampling temperature, 0 to 2, default 1. Lower is more deterministic — better for extraction and classification. |
| max_tokens | integer | Optional | Maximum tokens to generate. If omitted, it follows the model context length and the space left in it. |
curl https://api.yrinvestment.net/v1/chat/completions \ -H "Authorization: Bearer $YR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "DeepSeek-V4.1-Flash", "messages": [ {"role": "system", "content": "You are a precise technical assistant."}, {"role": "user", "content": "In one sentence, what problem does heterogeneous compute scheduling solve?"} ], "temperature": 0.6, "max_tokens": 512, "stream": false }'
import os from openai import OpenAI client = OpenAI( api_key=os.environ["YR_API_KEY"], base_url="https://api.yrinvestment.net/v1", # change only this line ) resp = client.chat.completions.create( model="DeepSeek-V4.1-Flash", messages=[ {"role": "system", "content": "You are a precise technical assistant."}, {"role": "user", "content": "In one sentence, what problem does heterogeneous compute scheduling solve?"}, ], temperature=0.6, max_tokens=512, ) print(resp.choices[0].message.content) print(resp.usage.total_tokens) print(resp.model_extra["yr_routing"]) # which silicon and which compute centre served this call
import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.YR_API_KEY, baseURL: "https://api.yrinvestment.net/v1", }); const resp = await client.chat.completions.create({ model: "Qwen3-32B", messages: [ { role: "system", content: "You are a precise technical assistant." }, { role: "user", content: "In one sentence, what problem does heterogeneous compute scheduling solve?" }, ], temperature: 0.6, max_tokens: 512, }); console.log(resp.choices[0].message.content); console.log(resp.usage.total_tokens); console.log(resp.yr_routing);
yr_routing. It is a platform extension on top of the standard response;
the standard SDKs drop it into their unknown-field container (model_extra in Python, a plain property in Node) and nothing about the official semantics changes.
Set stream to true and the endpoint switches to Server-Sent Events:
the response header is text/event-stream, the body is a series of data: lines, each one a JSON chunk,
and the text delta sits in choices[0].delta.content. The final chunk carries finish_reason and usage,
followed by data: [DONE].
data: {"id":"chatcmpl-8f2a1c7d4b9e0a3f5c6d","object":"chat.completion.chunk","created":1789243100,"model":"DeepSeek-V4.1-Flash","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}],"yr_routing":{"chip":"h200","chip_name":"H200","site":"HK-TKO","ttft_ms":180}}
data: {"id":"chatcmpl-8f2a1c7d4b9e0a3f5c6d","object":"chat.completion.chunk","created":1789243100,"model":"DeepSeek-V4.1-Flash","choices":[{"index":0,"delta":{"content":"Heterogene"},"finish_reason":null}]}
data: {"id":"chatcmpl-8f2a1c7d4b9e0a3f5c6d","object":"chat.completion.chunk","created":1789243100,"model":"DeepSeek-V4.1-Flash","choices":[{"index":0,"delta":{"content":"ous sched"},"finish_reason":null}]}
data: {"id":"chatcmpl-8f2a1c7d4b9e0a3f5c6d","object":"chat.completion.chunk","created":1789243100,"model":"DeepSeek-V4.1-Flash","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":42,"completion_tokens":168,"total_tokens":210}}
data: [DONE]
yr_routingA platform-specific extension field, present on the first streamed chunk (and at the top level of a non-streamed response). It tells you where the call actually landed — the silicon underneath is heterogeneous, but you never have to guess which card you got.
| Field | Example | Meaning |
|---|---|---|
| chip | h200 | Identifier of the silicon family that served this inference. |
| chip_name | H200 | Display name of that silicon family — H200, H100, Ascend A2, Hygon DCU. |
| site | HK-TKO | Compute centre code, useful for proximity routing and for checking data residency. |
| ttft_ms | 180 | Time to first token for this call, in milliseconds. |
If the workload has hard requirements on jurisdiction or silicon family — domestic cards only, data must not leave the country — set a routing policy on the relevant key in the console, then use this field to verify it client-side.
stream = client.chat.completions.create( model="DeepSeek-V4.1-Flash", messages=[{"role": "user", "content": "Hello"}], stream=True, ) for chunk in stream: routing = getattr(chunk, "yr_routing", None) if routing: print("served by", routing["chip_name"], "@", routing["site"]) print(chunk.choices[0].delta.content or "", end="")
GET /v1/modelsReturns every model currently live, in the same shape as the OpenAI models endpoint: object is list,
data is the model array, id is the model name you drop straight into the model field, and
owned_by is the vendor. Use it to pull the selectable models at application start instead of hard-coding them.
curl https://api.yrinvestment.net/v1/models \
-H "Authorization: Bearer $YR_API_KEY"
{
"object": "list",
"data": [
{ "id": "DeepSeek-V4.1-Flash", "object": "model", "owned_by": "DeepSeek", "created": 1756000000 },
{ "id": "Qwen3-32B", "object": "model", "owned_by": "Qwen", "created": 1756000000 }
]
}
Which silicon families each model runs on, and what the evidence is, is set out in the compatibility matrix.
Each API key carries three independent limits. They do not substitute for one another, and hitting any one of them rejects the request. The current level of all three is live in the console.
| Limit | Counted as | Default | On breach |
|---|---|---|---|
| RPM | Requests per minute, counted per key on a sliding window | 3,000 | Returns 429 — retry with exponential backoff |
| TPM | Tokens per minute, input and output combined | 8,000,000 | Returns 429 — split the batch or lower concurrency |
| Monthly token quota | Total tokens the key has consumed within the calendar month | Per key | Returns 401 — raise the quota in the console or move to another key |
Errors return JSON in a fixed shape: {"error": {"message": "...", "code": 401}}.
message is written to be shown to an operator as-is, and code matches the HTTP status.
| Status | Meaning | Typical cause and fix |
|---|---|---|
| 400 | Malformed request | Missing model or messages, a wrong field type, or unparseable JSON. Check the body and resend. |
| 401 | Auth failed or quota spent | Missing Authorization header, invalid key, disabled key, or the monthly token quota for that key is spent. |
| 403 | Forbidden | The account role cannot access this view, or the key is not allowed to call this model. |
| 404 | Endpoint or resource not found | Wrong path, or the requested model is not live. Check the name against /v1/models. |
| 429 | Rate limited | RPM or TPM exceeded. Retry with exponential backoff, lower concurrency, or split the batch. |
| 500 | Internal server error | A platform-side fault. Send us the request id from the response and we can pin it to the exact scheduling record. |