AI Proxy Multi Configuration
Static Configurations
APISIX 3.18.0 uses ngx_http_ffi_client by default for upstream requests from ai-proxy, ai-proxy-multi, and ai-request-rewrite. Set http_client to lua-resty-http to use the Lua client instead. API7 Gateway 3.9 and 3.10 use the Lua client and do not expose this setting.
To use the Lua client in an APISIX host or Docker deployment, configure the following setting:
plugin_attr:
ai-proxy:
http_client: lua-resty-httpThen reload APISIX for the change to take effect.
Export the full effective values for the installed APISIX release:
helm get values <release-name> -n <namespace> --all -o yaml > values.yamlAdd or update the following value:
apisix:
pluginAttrs:
ai-proxy:
http_client: lua-resty-httpThen apply the values file with the chart used for this APISIX release:
helm upgrade <release-name> apisix/apisix -n <namespace> -f values.yamlParameters
See plugin common configurations for configuration options available to all plugins.
-
fallback_strategy—string or array· optionalValid values: string:
instance_health_and_rate_limiting,http_429, orhttp_5xx
array: Any combination ofrate_limiting,http_429, andhttp_5xxFallback strategy. The option
instance_health_and_rate_limitingis kept for backward compatibility and is functionally the same asrate_limiting.With
rate_limitingorinstance_health_and_rate_limiting, when the current instance's quota is exhausted, the request is forwarded to the next instance regardless of priority. Withhttp_429, if an instance returns status code 429, the request is retried with other instances. Withhttp_5xx, if an instance returns a 5xx status code, the request is retried with other instances. If all instances fail, the plugin returns the last upstream status, response body, andContent-Type.When not set, the plugin will not forward the request to low priority instances when tokens of the high priority instance are exhausted.
-
max_retries—integer· optionalValid values: greater than or equal to 0
Maximum number of fallback retries after the initial request fails. This bounds how many additional instances a single request tries, so it does not exhaust every configured instance. Only takes effect together with
fallback_strategy. When not set, there is no explicit cap and the plugin retries until an instance succeeds or all instances have been tried. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. -
retry_on_failure_within_ms—integer· optionalValid values: greater than or equal to 1
Only fall back to another instance when the upstream fails within this many milliseconds. Fast failures (such as connection errors and quick 429 or 5xx responses) are retried, while a slow failure that takes longer than this is returned to the client directly to avoid doubling the total wait time. Only takes effect together with
fallback_strategy. When not set, the plugin retries regardless of how long the failed attempt took. Available in API7 Enterprise from version 3.9.14 and APISIX from version 3.17.0. -
fallback_http_statuses—array[integer]· optionalValid values: each between 400 and 599, no duplicates
Additional upstream HTTP status codes that make the request fall back to another instance, on top of the
http_429andhttp_5xxentries offallback_strategy. Use it for statuses that mean the instance's own credential or quota is the problem rather than the request, such as401for an expired key or403for a disabled account, so the request is retried elsewhere instead of being returned to the client. Only takes effect together withfallback_strategy. Available in API7 Enterprise from version 3.9.19 on the 3.9 line and from version 3.10.6 on the 3.10 line. -
balancer—object· optionalLoad balancing configurations.
-
algorithm—string· optional · default:roundrobinValid values:
roundrobin,chash, orsemanticLoad balancing algorithm. When set to
roundrobin, weighted round robin algorithm is used. When set tochash, consistent hashing algorithm is used. When set tosemantic, the instance whoseexamplesare semantically closest to the prompt is used, configured undersemantic_opts.The
semanticalgorithm does not participate in health checks,fallback_strategy, ormax_retries. An upstream failure on the selected instance is returned to the client; the algorithm falls back only when no instance clears its threshold or the embedding request fails.The
semanticalgorithm is available in API7 Enterprise from version 3.9.18 on the 3.9 line and from version 3.10.5 on the 3.10 line, and in APISIX from version 3.18.0. -
hash_on—string· optional · default:varsValid values:
vars,header,cookie,consumer, orvars_combinationsUsed when
typeischash. Support hashing on built-in variables, header, cookie, consumer, or a combination of built-in variables. -
key—string· optionalUsed when
typeischash. Whenhash_onis set toheaderorcookie,keyis required. Whenhash_onis set toconsumer,keyis not required as the consumer name will be used as the key automatically.
-
-
semantic_opts—object· optionalConfigurations for the
semanticbalancer algorithm. Required whenbalancer.algorithmissemantic, and ignored otherwise.Available in API7 Enterprise from version 3.9.18 on the 3.9 line and from version 3.10.5 on the 3.10 line, and in APISIX from version 3.18.0.
-
embeddings—object· requiredEmbedding service used to turn the prompt and each instance's
examplesinto vectors. The prompt is embedded on every request, so this service is on the request path.-
provider—string· requiredValid values:
openaiorazure-openaiEmbedding service provider.
-
model—string· requiredName of the embedding model, such as
text-embedding-3-small. -
endpoint—string· optionalEmbedding API endpoint. Optional for
openai, which defaults to the public API. Required forazure-openai, where it has to be the full URL, such ashttps://{resource}.openai.azure.com/openai/deployments/{deployment}/embeddings?api-version={version}. -
auth—object· requiredAuthentication for the embedding service, carried either as headers or as query parameters.
-
header—object· optionalKey-value pairs sent as request headers to the embedding service.
-
query—object· optionalKey-value pairs sent as query parameters to the embedding service.
-
-
timeout—integer· optional · default:3000Valid values: greater than or equal to 1
Timeout in milliseconds for an embedding request. Because the prompt is embedded synchronously, this bounds the latency added to each request when the embedding service is slow. On a timeout the request is routed to the fallback instance rather than failed.
-
ssl_verify—boolean· optional · default:trueIf true, verify the embedding service's TLS certificate.
-
-
threshold—number· optional · default:0Valid values: between -1 and 1 inclusive
Global minimum cosine similarity an instance has to reach to be selected. An instance's own
thresholdoverrides this value. The default of0admits almost any prompt, so the fallback instance is only ever reached once a threshold above0is set. -
fallback—string· optionalName of the instance to route to when no instance clears its threshold or the embedding request fails. It is otherwise a normally ranked instance and needs its own
examples. Defaults to the first instance when unset. -
debugging—boolean· optional · default:falseIf true, return the per-instance similarity scores and the routing decision in the
X-AI-Semantic-ScoresandX-AI-Semantic-Picked-Instanceresponse headers. Intended for tuningexamplesand thresholds, not for production traffic.
-
-
instances—array[object]· requiredLLM instance configurations.
-
name—string· requiredName of the LLM service instance.
-
examples—array[string]· optionalValid values: between 1 and 64 items
Example utterances representing what this instance handles. Each one is embedded into its own reference vector, and the semantic balancer routes a request to the instance whose closest example is most similar to the prompt.
Required for every instance when
balancer.algorithmissemantic, including the instance named bysemantic_opts.fallback. Ignored by the other algorithms.Available in API7 Enterprise from version 3.9.18 on the 3.9 line and from version 3.10.5 on the 3.10 line, and in APISIX from version 3.18.0.
-
threshold—number· optionalValid values: between -1 and 1 inclusive
Minimum cosine similarity a prompt has to reach for this instance to be selected by the semantic balancer. Overrides
semantic_opts.thresholdfor this instance.Available in API7 Enterprise from version 3.9.18 on the 3.9 line and from version 3.10.5 on the 3.10 line, and in APISIX from version 3.18.0.
-
provider—string· requiredValid values:
openai,deepseek,azure-openai,aimlapi,gemini,vertex-ai,anthropic,openrouter,bedrock,openai-compatibleLLM service provider.
When set to
openai, the plugin sends detected Chat Completions, Responses API, and Embeddings requests to their corresponding OpenAI endpoints.When set to
deepseek, the plugin will proxy requests tohttps://api.deepseek.com/chat/completions.When set to
gemini(available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests tohttps://generativelanguage.googleapis.com/v1beta/openai/chat/completions. If you are proxying requests to an embedding model, you should configure the embedding model endpoint in theoverride.When set to
vertex-ai(available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin proxies requests to Google Cloud Vertex AI. For chat completions, the plugin will proxy requests tohttps://{region}-aiplatform.googleapis.com/v1beta1/projects/{project_id}/locations/{region}/endpoints/openapi/chat/completions. For embeddings, the plugin will proxy requests tohttps://{region}-aiplatform.googleapis.com/v1/projects/{project_id}/locations/{region}/publishers/google/models/{model}:predict. These require configuringprovider_confwithproject_idandregion. Alternatively, you can configureoverridefor a custom endpoint.When set to
anthropic(available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin sends detected Chat Completions requests tohttps://api.anthropic.com/v1/chat/completionsand native Anthropic Messages requests tohttps://api.anthropic.com/v1/messages.When set to
openrouter(available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests tohttps://openrouter.ai/api/v1/chat/completions.When set to
bedrock(available from API7 Enterprise 3.9.12 and APISIX 3.17.0), the plugin proxies requests to AWS Bedrock using the Converse API.When set to
aimlapi(available from APISIX 3.14.0 and Enterprise 3.8.17), the plugin uses the OpenAI-compatible driver and proxies the request tohttps://api.aimlapi.com/v1/chat/completions.When set to
openai-compatible, the plugin proxies requests to the custom endpoint configured inoverride.When set to
azure-openai, the plugin also proxies requests to the custom endpoint configured inoverrideand additionally removes themodelparameter from user requests. -
priority—integer· optional · default:0Priority of the LLM instance in load balancing.
prioritytakes precedence overweight. -
weight—integer· requiredValid values: greater than or equal to 0
Weight of the LLM instance in load balancing.
-
auth—object· requiredAuthentication configurations.
-
header—object· optionalAuthentication headers. At least one of the
headerandqueryshould be configured. You can configure additional custom headers that will be forwarded to the upstream LLM service. -
query—object· optionalAuthentication query parameters. At least one of the
headerandqueryshould be configured. -
gcp—object· optionalGCP service account authentication for Vertex AI. Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0.
-
service_account_json—string· optionalGCP service account JSON content used for authentication. This can be configured using this parameter or by setting the
GCP_SERVICE_ACCOUNTenvironment variable. -
max_ttl—integer· optionalMaximum TTL for GCP access token caching, in seconds.
-
expire_early_secs—integer· optional · default:60Number of seconds to expire the access token before its actual expiration time. This prevents edge cases where tokens expire during active requests.
-
-
aws—object· optionalAWS IAM credentials for SigV4 signing. Required when
providerisbedrock(for Bedrock,auth.awsis sufficient andauth.header/auth.queryare not required). Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0.-
access_key_id—string· requiredAWS IAM access key ID.
-
secret_access_key—string· requiredAWS IAM secret access key.
-
session_token—string· optionalAWS session token for temporary credentials (e.g. from STS AssumeRole).
-
-
-
options—object· optionalModel configurations.
In addition to
model, you can configure additional parameters and they will be forwarded to the upstream LLM service in the request body. For instance, if you are working with OpenAI or DeepSeek, you can configure additional parameters such asmax_tokens,temperature,top_p, andstream. See your LLM provider's API documentation for more available options.-
model—string· optionalName of the LLM model, such as
gpt-4orgpt-3.5. See your LLM provider's API documentation for more available models.
-
-
provider_conf—object· optionalProvider-specific configuration. Required when
providerisbedrock. Whenproviderisvertex-ai, configure eitherprovider_conforoverride.endpoint.Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0.
-
project_id—string· optionalGoogle Cloud Project ID for Vertex AI.
-
region—string· requiredCloud region. For
vertex-ai, this is the GCP region. Forbedrock, this is the AWS region (e.g.us-east-1).
-
-
override—object· optionalOverride setting.
-
endpoint—string· optionalLLM provider endpoint to replace the endpoint selected for the detected request protocol.
-
llm_options—object· optionalProvider-aware LLM option overrides. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.
-
max_tokens—integer· optionalMaximum number of output tokens. The gateway automatically maps this to the correct field name for the target provider, such as
max_completion_tokensfor OpenAI Chat ormax_output_tokensfor OpenAI Responses API, and overwrites the client value.
-
-
request_body—object· optionalPer target-protocol request body overrides. Keys are target protocol names, such as
openai-chat,openai-responses,openai-embeddings,anthropic-messages,bedrock-converse, andpassthrough. Values are partial request bodies that are deep-merged into the outgoing body. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. -
request_body_force_override—boolean· optional · default:falseWhen
false(default), client request body fields take priority andrequest_bodyoverride values only fill in missing fields. Whentrue,request_bodyoverride values overwrite client fields. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.
-
-
checks—object· optionalHealth check configurations.
Note that at the moment, OpenAI and DeepSeek do not provide an official health check endpoint. Other LLM services that you can configure under
openai-compatibleprovider may have available health check endpoints.-
active—object· requiredActive health check configurations.
-
type—string· optional · default:httpValid values:
http,https, ortcpType of health check connection.
-
timeout—number· optional · default:1Health check timeout in seconds.
-
concurrency—integer· optional · default:10Number of upstream nodes to be checked at the same time.
-
host—string· optionalHTTP host.
-
port—integer· optionalValid values: between 1 and 65535 inclusive
HTTP port.
-
http_path—string· optional · default:/Path for HTTP probing requests.
-
http_method—string· optional · default:GETValid values:
CONNECT,DELETE,GET,HEAD,OPTIONS,PATCH,POST,PURGE,PUT, orTRACEHTTP method for active health check probing requests. Available in API7 Enterprise and in APISIX from version 3.18.0.
-
http_req_body—string· optionalRequest body to send in active health check probing requests. This is useful when
http_methodis set toPOST. Defaults to empty string. Available in API7 Enterprise and in APISIX from version 3.18.0. -
https_verify_certificate—boolean· optional · default:trueIf true, verify the node's TLS certificate.
-
healthy—object· optionalHealthy check configurations.
-
interval—integer· optional · default:1Time interval of checking healthy nodes, in seconds.
-
http_statuses—array[integer]· optional · default:[200,302]Valid values: status code between 200 and 599 inclusive
An array of HTTP status codes that defines a healthy node.
-
successes—integer· optional · default:2Valid values: between 1 and 254 inclusive
Number of successful probes to define a healthy node.
-
-
req_headers—array[string]· optionalList of additional HTTP headers to send in health check probing requests, in
"Header: Value"format. -
unhealthy—object· optionalUnhealthy check configurations.
-
interval—integer· optional · default:1Time interval of checking unhealthy nodes, in seconds.
-
http_statuses—array[integer]· optional · default:[429,404,500,501,502,503,504,505]Valid values: status code between 200 and 599 inclusive
An array of HTTP status codes that defines an unhealthy node.
-
http_failures—integer· optional · default:5Valid values: between 1 and 254 inclusive
Number of HTTP failures to define an unhealthy node.
-
tcp_failures—integer· optional · default:2Valid values: between 1 and 254 inclusive
Number of TCP failures to define an unhealthy node.
-
timeouts—integer· optional · default:3Valid values: between 1 and 254 inclusive
Number of probe timeouts to define an unhealthy node.
-
-
-
-
-
logging—object· optionalLogging configurations. These configurations apply to access logs and logs sent to logging plugins, and do not affect the error log.
-
summaries—boolean· optional · default:falseIf true, add an
llm_summaryobject to logger entries with model, latency, and token usage. In API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0, the summary also includes stream status, tool count and usage, end-user ID, cache read and creation tokens, reasoning tokens, and content risk level when available. -
payloads—boolean· optional · default:falseIf true, log request and response payload.
-
-
timeout—integer· optional · default:30000Valid values: between 1 and 600000 inclusive
Timeout in milliseconds for each connect, send, or blocking read operation to the LLM service. It does not limit the total duration of a streaming response; use
max_stream_duration_msfor that limit. -
max_req_body_size—integer· optional · default:67108864Valid values: greater than or equal to 1
Maximum request body size in bytes that the plugin reads into memory. Larger requests are rejected with HTTP 413. This prevents unbounded memory buffering of large request bodies. The default is 67108864 bytes (64 MiB). Available in API7 Enterprise from versions 3.9.14 and 3.10.1 in their respective release lines, and APISIX from version 3.17.0.
-
max_stream_duration_ms—integer· optionalValid values: greater than or equal to 1
Maximum wall-clock duration, in milliseconds, for a streaming AI response. The limit is optional. When reached, the gateway closes the connection; if output has already started, the stream ends without a protocol terminator such as
[DONE],message_stop, orresponse.completed. Enforcement occurs between upstream reads, so the final chunk can exceed the configured duration. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. -
max_response_bytes—integer· optionalValid values: greater than or equal to 1
Maximum total bytes read from the upstream for one streaming or non-streaming AI response. The limit is optional and checked between upstream reads, so the final chunk can exceed it. If the limit is exceeded before output starts, the gateway returns
502 Bad Gateway; after output starts, the gateway closes the stream without a protocol terminator. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0. -
streaming_flush_interval_ms—integer· optional · default:10Valid values: greater than or equal to 0
Background flush interval in milliseconds for streaming responses. A positive value periodically flushes buffered output to bound client latency when the upstream sends tokens in bursts. Set to 0 to flush each chunk synchronously. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0.
-
keepalive—boolean· optional · default:trueIf true, keep the connection alive when requesting the LLM service.
-
keepalive_timeout—integer· optional · default:60000Valid values: greater than or equal to 1000
Keepalive timeout in milliseconds when requesting the LLM service.
-
keepalive_pool—integer· optional · default:30Valid values: greater than or equal to 1
Keepalive pool size for when connecting with the LLM service.
-
ssl_verify—boolean· optional · default:trueIf true, verify the LLM service's certificate.