API7 Docs

AI Proxy Configuration

Static Configurations

APISIX 3.18.0 uses ngx_http_ffi_client by default for upstream requests from ai-proxy, ai-proxy-multi, and ai-request-rewrite. Set http_client to lua-resty-http to use the Lua client instead. API7 Gateway 3.9 and 3.10 use the Lua client and do not expose this setting.

To use the Lua client in an APISIX host or Docker deployment, configure the following setting:

config.yaml
plugin_attr:
  ai-proxy:
    http_client: lua-resty-http

Then reload APISIX for the change to take effect.

Parameters

See plugin common configurations for configuration options available to all plugins.

  • providerstring · required

    Valid values: openai, deepseek, azure-openai, aimlapi, gemini, vertex-ai, anthropic, openrouter, bedrock, openai-compatible

    LLM service provider.

    When set to openai, the plugin sends detected Chat Completions, Responses API, and Embeddings requests to their corresponding OpenAI endpoints.

    When set to deepseek, the plugin will proxy requests to https://api.deepseek.com/chat/completions.

    When set to gemini (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to https://generativelanguage.googleapis.com/v1beta/openai/chat/completions. If you are proxying requests to an embedding model, you should configure the embedding model endpoint in the override.

    When set to vertex-ai (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin proxies requests to Google Cloud Vertex AI. For chat completions, the plugin will proxy requests to https://{region}-aiplatform.googleapis.com/v1beta1/projects/{project_id}/locations/{region}/endpoints/openapi/chat/completions. For embeddings, the plugin will proxy requests to https://{region}-aiplatform.googleapis.com/v1/projects/{project_id}/locations/{region}/publishers/google/models/{model}:predict. These require configuring provider_conf with project_id and region. Alternatively, you can configure override for a custom endpoint.

    When set to anthropic (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin sends detected Chat Completions requests to https://api.anthropic.com/v1/chat/completions and native Anthropic Messages requests to https://api.anthropic.com/v1/messages.

    When set to openrouter (available from APISIX 3.15.0 and Enterprise 3.9.2), the plugin will proxy requests to https://openrouter.ai/api/v1/chat/completions.

    When set to bedrock (available from API7 Enterprise 3.9.12 and APISIX 3.17.0), the plugin proxies requests to AWS Bedrock using the Converse API. Requires configuring auth.aws with IAM credentials and provider_conf.region with the AWS region. Supports both non-streaming and streaming (ConverseStream) when stream is set to true in the request body.

    When set to aimlapi (available from APISIX 3.14.0 and Enterprise 3.8.17), the plugin uses the OpenAI-compatible driver and proxies the request to https://api.aimlapi.com/v1/chat/completions.

    When set to openai-compatible, the plugin proxies requests to the custom endpoint configured in override.

    When set to azure-openai, the plugin also proxies requests to the custom endpoint configured in override and additionally removes the model parameter from user requests.

  • authobject · required

    Authentication configurations.

    • headerobject · optional

      Authentication headers.

    • queryobject · optional

      Authentication query parameters.

    • gcpobject · optional

      GCP service account authentication for Vertex AI. Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0.

      • service_account_jsonstring · optional

        GCP service account JSON content used for authentication. This can be configured using this parameter or by setting the GCP_SERVICE_ACCOUNT environment variable.

      • max_ttlinteger · optional

        Maximum TTL for GCP access token caching, in seconds.

      • expire_early_secsinteger · optional · default: 60

        Number of seconds to expire the access token before its actual expiration time. This prevents edge cases where tokens expire during active requests.

    • awsobject · optional

      AWS IAM credentials for SigV4 signing. Required when provider is bedrock (for Bedrock, auth.aws is sufficient and auth.header/auth.query are not required). Available in API7 Enterprise from version 3.9.12 and APISIX from version 3.17.0.

      • access_key_idstring · required

        AWS IAM access key ID.

      • secret_access_keystring · required

        AWS IAM secret access key.

      • session_tokenstring · optional

        AWS session token for temporary credentials (e.g. from STS AssumeRole).

  • optionsobject · optional

    Model configurations.

    In addition to model, you can configure additional parameters and they will be forwarded to the upstream LLM service in the request body. For instance, if you are working with OpenAI, you can configure additional parameters such as temperature, top_p, and stream. See your LLM provider's API documentation for more available options.

    • modelstring · optional

      Name of the LLM model, such as gpt-4 or gpt-3.5. See your LLM provider's API documentation for more available models.

  • provider_confobject · optional

    Provider-specific configuration. Required when provider is bedrock. When provider is vertex-ai, configure either provider_conf or override.endpoint.

    Available in API7 Enterprise from 3.9.2 and APISIX from version 3.17.0.

    • project_idstring · optional

      Google Cloud Project ID. Required when provider is vertex-ai.

    • regionstring · required

      Cloud region. For vertex-ai, this is the GCP region. For bedrock, this is the AWS region (e.g. us-east-1).

  • overrideobject · optional

    Override setting.

    • endpointstring · optional

      LLM provider endpoint. Required when provider is openai-compatible.

    • llm_optionsobject · optional

      Provider-aware LLM option overrides. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

      • max_tokensinteger · optional

        Maximum number of output tokens. The gateway automatically maps this to the correct field name for the target provider (e.g. max_completion_tokens for OpenAI Chat, max_output_tokens for OpenAI Responses API). Always force-overwrites the client value.

    • request_bodyobject · optional

      Per target-protocol request body overrides. Keys are target protocol names (openai-chat, openai-responses, openai-embeddings, anthropic-messages, bedrock-converse, passthrough); values are partial request bodies that are deep-merged into the outgoing body (objects merged recursively, arrays and scalars replaced wholesale). Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

    • request_body_force_overrideboolean · optional · default: false

      When false (default), client request body fields take priority and request_body override values only fill in missing fields. When true, request_body override values forcefully overwrite client fields. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

  • loggingobject · optional

    Logging configurations. These configurations apply to access logs and logs sent to logging plugins, and do not affect the error log.

    • summariesboolean · optional · default: false

      If true, add an llm_summary object to logger entries with model, latency, and token usage. In API7 Enterprise 3.9.18 and 3.10.5, and in APISIX 3.18.0, the summary also includes stream status, tool count and usage, end-user ID, cache read and creation tokens, reasoning tokens, and content risk level when available.

    • payloadsboolean · optional · default: false

      If true, log request and response payload.

  • timeoutinteger · optional · default: 30000

    Valid values: between 1 and 600000 inclusive

    Timeout in milliseconds for each connect, send, or blocking read operation to the LLM service. It does not limit the total duration of a streaming response; use max_stream_duration_ms for that limit.

  • max_req_body_sizeinteger · optional · default: 67108864

    Maximum request body size in bytes that the plugin reads into memory (default 67108864 bytes, which is 64 MiB). Requests with a body larger than this limit are rejected with HTTP 413. This prevents unbounded memory buffering of large request bodies. Available in API7 Enterprise from versions 3.9.14 and 3.10.1 in their respective release lines, and APISIX from version 3.17.0.

  • keepaliveboolean · optional · default: true

    If true, keep the connection alive when requesting the LLM service.

  • keepalive_timeoutinteger · optional · default: 60000

    Valid values: greater than or equal to 1000

    Keepalive timeout in milliseconds when requesting the LLM service.

  • keepalive_poolinteger · optional · default: 30

    Valid values: greater than or equal to 1

    Keepalive pool size for when connecting with the LLM service.

  • ssl_verifyboolean · optional · default: true

    If true, verify the LLM service's certificate.

  • max_stream_duration_msinteger · optional

    Maximum wall-clock duration, in milliseconds, for a streaming AI response. The limit is optional. When reached, the gateway closes the connection; if output has already started, the stream ends without a protocol terminator such as [DONE], message_stop, or response.completed. Enforcement occurs between upstream reads, so the final chunk can exceed the configured duration. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

  • max_response_bytesinteger · optional

    Maximum total bytes read from the upstream for one streaming or non-streaming AI response. The limit is optional and checked between upstream reads, so the final chunk can exceed it. If the limit is exceeded before output starts, the gateway returns 502 Bad Gateway; after output starts, the gateway closes the stream without a protocol terminator. Available in API7 Enterprise from version 3.9.10 and APISIX from version 3.17.0.

  • streaming_flush_interval_msinteger · optional · default: 10

    Background flush interval in milliseconds for streaming responses. A positive value starts a background thread that flushes output periodically to bound client latency when upstreams burst multiple tokens at once. Set to 0 to flush each chunk synchronously inline. Available in API7 Enterprise from version 3.9.13 and APISIX from version 3.17.0.