Model serving

Model Catalog

One API key for every model — language, multimodal, embedding, reranking and speech. OpenAI-compatible, with each call routed to silicon that can actually run it, according to the compatibility matrix.

—
Models live
—
Vendors
—
Modalities
Category
Vendor
Silicon family
Search
Loading
Compatibility matrix

Which models run on which silicon

The table lists compatibility per model per silicon family, and the wording of each cell reflects how strong the evidence is. Officially supported means the vendor publishes it. Experimental means it runs but is not on the official list. Community only means the sole evidence is community reports. Not verified means we found no first-hand source to cite.

Model nameScaleArchitecture NVIDIAAscend A2·A3Hygon DCUCambriconMetaX
The MetaX column, and nearly all of the Cambricon column, read Not verified. That is the honest state of the evidence, not an omission. Both vendors publish only a vLLM plugin version, not a per-model support list. A plugin is not a model list: for a specific model, either the vendor issues written confirmation or we measure it ourselves. The Hygon column currently rests on community write-ups and press coverage with no first-hand vendor source, and is labelled accordingly. The Ascend column follows the official vLLM-Ascend support matrix.

Get an API key the moment you register

One key calls every model above. Quota, rate limits and usage are live in the console.

Get started → Read the docs