API7 Docs
AISIX AI GatewayOpen-Source AISIX Gateway Admin APIGet Model by ID

Get Model by ID

GET /admin/v1/models/{id}

GET /admin/v1/models/{id}

Interactive request editor loads with JavaScript.

Authorization

AdminBearer
AuthorizationBearer <token>

Admin key from config.admin.admin_keys.

In: header

Path Parameters

id*string

Model resource ID, as assigned by the active resource source (a UUIDv5 derived from the entry name in file mode; the etcd key's ID segment otherwise).

Response Body

application/json

application/json

application/json

application/json

curl -X GET "http://127.0.0.1:3001/admin/v1/models/c7f13a2e-5d6b-4b87-9f47-2c34de35f4aa"

OK

{
  "id": "string",
  "revision": 0,
  "value": {
    "allowed_cidrs": [
      "string"
    ],
    "auto_prompt_caching": {
      "enabled": true,
      "ttl": "5m"
    },
    "background_model_check": {
      "enabled": true,
      "ignore_statuses": [
        100
      ],
      "interval_seconds": 5,
      "max_tokens": 1,
      "prompt": "string",
      "stale_after_seconds": 1,
      "timeout_seconds": 1
    },
    "cooldown": {
      "default_seconds": 30,
      "enabled": false,
      "honor_retry_after": true,
      "max_seconds": 600,
      "trigger_on_timeout": true,
      "trigger_on_transport": true,
      "trigger_statuses": [
        401,
        408,
        429,
        500,
        502,
        503,
        504
      ]
    },
    "cost": {
      "input_per_1k": 0,
      "output_per_1k": 0
    },
    "display_name": "string",
    "effort_mapping": {
      "property1": "string",
      "property2": "string"
    },
    "embedding": {
      "dimensions": 1,
      "normalize": true
    },
    "ensemble": {
      "judge": {
        "model": "string",
        "synthesis_prompt": "string"
      },
      "min_responses": 1,
      "panel": [
        {
          "model": "string",
          "seed": 0,
          "temperature": 0,
          "weight": 0
        }
      ],
      "timeout_ms": 0
    },
    "model_name": "string",
    "provider": "string",
    "provider_key_id": "string",
    "rate_limit": {
      "concurrency": 0,
      "rpd": 0,
      "rph": 0,
      "rpm": 0,
      "rps": 0,
      "tpd": 0,
      "tpm": 0
    },
    "retries": 0,
    "routing": {
      "fallback_on_statuses": [
        400
      ],
      "hash_on": [
        {
          "name": "string",
          "type": "header"
        }
      ],
      "max_fallbacks": 0,
      "retries": 0,
      "retry_on_429": true,
      "strategy": "failover",
      "targets": [
        {
          "model": "string",
          "priority": 0,
          "tags": [
            "string"
          ],
          "weight": 0
        }
      ],
      "when_all_unavailable": "fail"
    },
    "semantic": {
      "default": "string",
      "embedding_model": "string",
      "embedding_timeout_ms": 0,
      "match": {
        "aggregation": "max",
        "distance_metric": "cosine",
        "threshold": 0
      },
      "on_embedding_failure": "default",
      "routes": [
        {
          "description": "string",
          "examples": [
            "string"
          ],
          "name": "string",
          "target": "string",
          "threshold": 0
        }
      ]
    },
    "stream_timeout": 0,
    "timeout": 0
  }
}
Complete operation details and schema variants

GET /admin/v1/models/{id}

Get Model by ID

Get a model resource by ID.

Parameters

  • id (path, string, required): Model resource ID, as assigned by the active resource source (a UUIDv5 derived from the entry name in file mode; the etcd key's ID segment otherwise).

Responses

  • 200: OK
  • id (string, required): Resource ID, as assigned by the active resource source.
  • revision (integer, required): Monotonic resource revision: the etcd mod_revision of the entry, or the load generation in file mode.
  • value (object, required): Stored model configuration.
  • allOf variant 1
  • value.allowed_cidrs (array, optional): Client IP allowlist in CIDR notation. Empty or absent allows all clients.
  • value.auto_prompt_caching (object, optional): Automatic prompt-cache breakpoint injection for direct Anthropic models. Omit to leave injection off.
  • allOf variant 1
  • value.auto_prompt_caching.enabled (boolean, required): Whether automatic prompt-cache injection is active for this model.
  • value.auto_prompt_caching.ttl (object, optional): Cache lifetime for injected breakpoints: 5m (default when omitted) or 1h. A 1h cache write costs 2x the base input rate versus 1.25x for 5m, so it pays off only when the cached prefix is reused across a longer session.
  • allOf variant 1
  • value.background_model_check (object, optional): Direct-model-only background health-check configuration.
  • allOf variant 1
  • value.background_model_check.enabled (boolean, required): Whether background health checks are enabled for this model.
  • value.background_model_check.ignore_statuses (array, optional): Upstream status codes to ignore when evaluating background check failures.
  • value.background_model_check.interval_seconds (integer, required): Seconds between background health checks. Minimum: 5.
  • value.background_model_check.max_tokens (integer, required): Maximum completion tokens requested during each background health check.
  • value.background_model_check.prompt (string, required): Prompt sent to the model during each background health check.
  • value.background_model_check.stale_after_seconds (integer, required): Seconds after which the last completed background check is considered stale.
  • value.background_model_check.timeout_seconds (integer, required): Request timeout in seconds for each background health check. Minimum: 1.
  • value.cooldown (object, optional): Direct-model-only request-path cooldown configuration. Cooldown is opt-in: omit this field, or leave enabled unset, and the model is never taken out of rotation by request-path failures.
  • allOf variant 1
  • value.cooldown.default_seconds (integer, optional): Cooldown TTL in seconds when the upstream did not supply a Retry-After header or honor_retry_after is false.
  • value.cooldown.enabled (boolean, optional): Whether cooldown is active for this model. Cooldown is off unless this is set to true, so a model that omits it stays in rotation regardless of upstream failures.
  • value.cooldown.honor_retry_after (boolean, optional): Whether to use the upstream's Retry-After header as the cooldown TTL when it contains seconds.
  • value.cooldown.max_seconds (integer, optional): Upper bound on cooldown TTL when Retry-After is used.
  • value.cooldown.trigger_on_timeout (boolean, optional): Whether request-path timeouts trigger cooldown.
  • value.cooldown.trigger_on_transport (boolean, optional): Whether transport, decode, or stream-abort errors trigger cooldown.
  • value.cooldown.trigger_statuses (array, optional): Status codes that trigger cooldown, covering authentication failures, rate limits, and transient server errors. Caller-side validation errors such as 400, 403, and 422 are excluded.
  • value.cost (object, optional): Per-token cost for budget tracking. Omit it when cost tracking is not needed.
  • allOf variant 1
  • value.cost.input_per_1k (number, required): Prompt token cost in USD per 1,000 tokens.
  • value.cost.output_per_1k (number, required): Completion token cost in USD per 1,000 tokens.
  • value.display_name (string, required): Operator-facing unique label. Surfaces on /v1/models, req.model on chat completions, ApiKey.allowed_models, and the dashboard model list. Resource::name() returns this.
  • value.effort_mapping (object, optional): Direct-model-only mapping from a client-requested reasoning effort to the value sent upstream. The gateway applies one exact lookup after resolving the final target; unlisted values pass through unchanged.
  • value.embedding (object, optional): Embedding-modality metadata. Present on direct Models that serve an OpenAI-compatible /v1/embeddings endpoint (and can be referenced by a semantic router's embedding_model).
  • allOf variant 1
  • value.embedding.dimensions (integer, required): Output vector dimensionality. Used to validate vectors, key the example-vector cache, and (for endpoints that support it) request a reduced output size.
  • value.embedding.normalize (boolean, optional): Whether the endpoint already returns L2-normalized vectors. When false, the gateway normalizes before computing cosine similarity. Defaults to true.
  • value.ensemble (object, optional): Ensemble configuration for panel calls and judge synthesis.
  • allOf variant 1
  • value.ensemble.judge (object, required): Direct model that combines successful panel responses.
  • allOf variant 1
  • value.ensemble.judge.model (string, required): Model alias for the direct model that synthesizes panel responses.
  • value.ensemble.judge.synthesis_prompt (string, optional): Override for the built-in synthesis prompt template.
  • value.ensemble.min_responses (integer, optional): Minimum successful panel responses required before judge synthesis. When omitted, the gateway requires the smaller of 2 and the panel size.
  • value.ensemble.panel (array, required): Direct models called concurrently for each ensemble request.
  • value.ensemble.panel[].model (string, required): Model alias for a direct model that receives one panel request.
  • value.ensemble.panel[].seed (integer, optional): Sampling seed for this panel member.
  • value.ensemble.panel[].temperature (number, optional): Sampling temperature for this panel member. Omit it to keep the request's temperature.
  • value.ensemble.panel[].weight (integer, optional): Reserved for a future voting/quorum strategy. AISIX currently ignores this field.
  • value.ensemble.timeout_ms (integer, optional): Per-call upstream deadline applied to each panel member and the judge. Set 0 or omit it to disable the ensemble-level deadline.
  • value.model_name (string, optional): Upstream model identifier sent in provider requests. Routing and ensemble models leave this field unset.
  • value.provider (string, optional): Upstream vendor identity used for dispatch, compatibility checks, telemetry, and access logs. Routing and ensemble models leave this field unset.
  • value.provider_key_id (string, optional): Provider key resource ID used to authenticate upstream requests. Routing and ensemble models leave this field unset.
  • value.rate_limit (object, optional): Request, token, and concurrency limits for this model.
  • allOf variant 1
  • value.rate_limit.concurrency (integer,null, optional): Max concurrent in-flight requests.
  • value.rate_limit.rpd (integer,null, optional): Requests per 86,400-second window.
  • value.rate_limit.rph (integer,null, optional): Requests per 3,600-second window. There is no per-hour token limit field.
  • value.rate_limit.rpm (integer,null, optional): Requests per 60-second window.
  • value.rate_limit.rps (integer,null, optional): Requests per 1-second window. There is no per-second token limit field.
  • value.rate_limit.tpd (integer,null, optional): Tokens per 86,400-second window.
  • value.rate_limit.tpm (integer,null, optional): Tokens per 60-second window.
  • value.retries (integer, optional): Retry attempts against this model after a retryable upstream failure, before the request gives up (or, inside a model group, fails over to the next target). Absent falls back to the group's routing.retries, then to the deployment-wide upstream.retries default.
  • value.routing (object, optional): Virtual routing configuration. When set, the gateway selects a target from routing.targets and uses that target model's provider, model_name, and provider_key_id fields for upstream dispatch.
  • allOf variant 1
  • value.routing.fallback_on_statuses (array,null, optional): Additional upstream HTTP status codes that participate in retries and failover. By default a non-429 4xx response is treated as a caller error and returned as-is; providers that use 4xx codes for transient conditions (model overload, queue full, quota exhaustion) can be listed here, for example [408, 409]. 5xx codes are already retryable, so listing them changes nothing. Authentication (401/403) and validation (400) codes should only be listed when the provider is known to use them for transient failures.
  • value.routing.hash_on (array,null, optional): Where the consistent_hash hash key comes from: an ordered chain of sources, the first non-empty value winning. Defaults to the x-aisix-routing-key request header, falling back to the caller's API key id. Only valid with strategy: consistent_hash.
  • value.routing.hash_on[].name (string, optional): The header or cookie name to read. Required for header and cookie sources; not accepted for api_key or client_ip.
  • value.routing.hash_on[].type (object, required): Which request attribute supplies the hash key.
  • allOf variant 1
  • oneOf variant 1: Header
  • oneOf variant 2: Cookie
  • oneOf variant 3: API key
  • oneOf variant 4: Client IP
  • value.routing.max_fallbacks (integer,null, optional): Max number of later targets to attempt after the initial target fails permanently. When omitted, all later targets may be attempted.
  • value.routing.retries (integer,null, optional): Retry attempts on the current target before failing over, applied to every target that does not set its own retries. Absent falls back to the deployment-wide upstream.retries default.
  • value.routing.retry_on_429 (boolean,null, optional): Whether upstream 429 participates in retries and failover.
  • value.routing.strategy (object, optional): Strategy used to select a target for each request.
  • allOf variant 1
  • oneOf variant 1: Round robin
  • oneOf variant 2: Consistent hash
  • oneOf variant 3: Failover
  • oneOf variant 4: Least cost
  • oneOf variant 5: Least latency
  • oneOf variant 6: Least busy
  • value.routing.targets (array, required): Ordered set of direct models available to this routing model.
  • value.routing.targets[].model (string, required): Model alias for a direct model that can receive routed traffic.
  • value.routing.targets[].priority (integer, optional): Priority tier, default 0; a higher value is preferred (the APISIX node-priority convention — give backup targets -1). The strategy orders targets within each tier; a lower tier is only tried when every higher-tier target failed or is unavailable.
  • value.routing.targets[].tags (array, optional): Tags for tag/metadata-conditional routing. When a request carries routing tags, only targets whose tags intersect the request's are eligible; a target tagged "default" is the fallback used when nothing matches and for untagged requests. Absent/empty means the target opts out of tag filtering (eligible only via the default fallback once any sibling target is tagged). The configured strategy then orders whatever set survives.
  • value.routing.targets[].weight (integer, optional): Target weight, default 1. Used by round_robin (rotation share), consistent_hash (share of the hash ring), and least_busy (in-flight divided by weight). failover, least_cost, and least_latency accept the field but do not use it.
  • value.routing.when_all_unavailable (object, optional): Policy to apply when every target is unavailable because of runtime health or cooldown state.
  • allOf variant 1
  • oneOf variant 1: Fail
  • oneOf variant 2: Try anyway
  • value.semantic (object, optional): Semantic-routing configuration. When set, the gateway embeds the request and dispatches to the route whose examples it matches best, using that route's target Model for upstream dispatch.
  • allOf variant 1
  • value.semantic.default (string, required): Direct model alias used when no route clears its threshold.
  • value.semantic.embedding_model (string, required): Alias of an embedding-modality Model used to embed the request and (at apply time) the route examples.
  • value.semantic.embedding_timeout_ms (integer, optional): Per-call deadline for the embedding request in milliseconds. 0 or absent disables the embedding-specific deadline.
  • value.semantic.match (object, required): Shared matching parameters (metric, aggregation, default threshold).
  • allOf variant 1
  • value.semantic.match.aggregation (object, optional): Per-example score aggregation. v1: max.
  • allOf variant 1
  • oneOf variant 1: Max
  • value.semantic.match.distance_metric (object, optional): Similarity metric. v1: cosine.
  • allOf variant 1
  • oneOf variant 1: Cosine
  • value.semantic.match.threshold (number, required): Default similarity threshold for routes that do not set their own threshold. Higher is stricter.
  • value.semantic.on_embedding_failure (object, optional): Behavior when the embedding call fails or times out. Defaults to routing to default.
  • allOf variant 1
  • anyOf variant 1: Default or fail policy
  • allOf variant 1
  • oneOf variant 1: Default
  • oneOf variant 2: Fail
  • anyOf variant 2: Target model
  • value.semantic.on_embedding_failure.target (string, required): Direct-model alias to route to when embedding fails.
  • value.semantic.routes (array, required): Routes evaluated for each request. At least one is required.
  • value.semantic.routes[].description (string, optional): Human-facing description. Documentation only — v1 matches on examples, not on this field.
  • value.semantic.routes[].examples (array, required): Example utterances that define this route. AISIX embeds each example when applying the configuration and caches the vector. A request is matched against these examples. At least one example is required.
  • value.semantic.routes[].name (string, required): Operator-facing route label. Surfaced in the x-aisix-route response header and access logs (e.g. prod-chat -> route:legal).
  • value.semantic.routes[].target (string, required): Direct model alias that receives traffic matching this route.
  • value.semantic.routes[].threshold (number, optional): Per-route similarity threshold. A request matches this route only when its aggregated score is >= this value. When omitted, the router-level threshold applies.
  • value.stream_timeout (integer, optional): Maximum gap in milliseconds between upstream streaming chunks. 0 or absent falls back to the group's stream_timeout, then to the model's (or group's) timeout, then to the deployment-wide upstream.stream_timeout_ms / timeout_ms defaults.
  • value.timeout (integer, optional): End-to-end timeout in milliseconds for non-streaming upstream calls. Absent falls back to the group's timeout, then to the deployment-wide upstream.timeout_ms default. 0 disables the non-streaming timeout for this model.
  • oneOf variant 1: Routing model
  • oneOf variant 2: Direct model
  • oneOf variant 3: Ensemble model
  • oneOf variant 4: Semantic router
  • 401: Missing or invalid admin key
  • error_msg (string, required): Human-readable error message.
  • 404: Resource not found Response fields (application/json): identical to GET /admin/v1/models/{id}, response 401 above.
  • 500: Configuration store operation failed Response fields (application/json): identical to GET /admin/v1/models/{id}, response 401 above.

cURL

curl -X GET 'http://127.0.0.1:3001/admin/v1/models/{id}'