AI Cache Configuration
Parameters
See plugin common configurations for configuration options available to all plugins.
This plugin supports referencing sensitive parameter values from environment variables using the env:// prefix, or from a secret manager, such as HashiCorp Vault’s KV secrets engine, using the secret:// prefix. For more information, see environment variables in plugin and secrets.
-
exact—object· optionalSettings for exact-match caching, where a response is reused only when the normalized request is identical to a previously cached one.
-
ttl—integer· optional · default:3600Valid values: greater than or equal to 1
Time-to-live in seconds for a cached entry.
-
-
cache_key—object· optionalSettings that control how the cache key is scoped.
-
share_across_routes—boolean· optional · default:falseIf true, the cache key does not include the route ID, so identical requests on different routes can share cached responses. If false, each route has its own cache scope.
-
include_consumer—boolean· optional · default:falseIf true, the consumer name is included in the cache key, so cached responses are isolated per consumer.
-
include_vars—array[string]· optional · default:[]Names of additional context variables to include in the cache key scope, so requests with different values of these variables do not share cached responses.
-
-
max_cache_body_size—integer· optional · default:1048576Valid values: greater than or equal to 0
Maximum size in bytes of a response body that will be cached. Larger responses are not cached.
-
cache_headers—boolean· optional · default:trueIf true, the plugin adds the
X-AI-Cache-Statusresponse header (andX-AI-Cache-Ageon a cache hit). Set to false to omit these headers. -
fail_mode—string· optional · default:skipValid values:
skip,warn, orerrorBehavior when the request cannot be cached because no AI instance was selected, for example when the route does not also configure
ai-proxyorai-proxy-multi. Withskip, the request is passed through unchecked. Withwarn, the request is passed through and a warning is logged. Witherror, the request is rejected with HTTP 500. -
bypass_on—array[object]· optionalA list of request-header matching rules. If a request matches any rule, the cache is bypassed and the response is marked with
X-AI-Cache-Statusset toBYPASS.-
header—string· requiredValid values: non-empty
Name of the request header to match.
-
equals—string· requiredValue that the header must equal for the rule to match.
-
-
policy—string· optional · default:redisValid values:
redisCache storage backend. Currently only
redisis supported. -
layers—array[string]· optional · default:["exact"]Valid values:
exact, or bothexactandsemanticCache layers to enable.
exactis always required. Addsemanticto look for similar prompts through embeddings and vector search after an exact miss.Semantic caching was introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0.
-
semantic—object· optionalSettings for semantic caching. Required when
layersincludessemantic.Introduced in API7 Enterprise 3.9.16 and 3.10.3, and APISIX 3.18.0.
-
similarity_threshold—number· optional · default:0.95Valid values: between 0 and 1 inclusive
Minimum similarity score required for a semantic cache hit.
-
top_k—integer· optional · default:1Valid values: greater than or equal to 1
Number of nearest vector matches to retrieve.
-
distance_metric—string· optional · default:cosineValid values:
cosineDistance metric used by vector search.
-
ttl—integer· optional · default:86400Valid values: greater than or equal to 1
Time-to-live in seconds for semantic cache entries.
-
match—object· optionalSettings that control which conversation content is embedded for semantic matching.
-
message_countback—integer· optional · default:1Valid values: greater than or equal to 1
Number of recent user-message turns to include in the embedding input.
-
ignore_system_prompts—boolean· optional · default:trueIf true, system prompts are excluded from the embedding input.
-
ignore_assistant_prompts—boolean· optional · default:trueIf true, assistant messages are excluded from the embedding input.
-
ignore_tool_prompts—boolean· optional · default:trueIf true, tool messages are excluded from the embedding input.
-
-
embedding—object· requiredEmbedding provider configuration. Configure exactly one of
openaiorazure_openai.-
openai—object· optionalOpenAI-compatible embedding provider settings. Requires
modelandapi_key.-
endpoint—string· optionalOpenAI-compatible embedding API endpoint. When not configured, the public OpenAI embeddings endpoint is used.
-
model—string· requiredEmbedding model name, such as
text-embedding-3-small. -
api_key—string· requiredAPI key used to authenticate with the embedding provider. The value is encrypted with AES before being stored in etcd.
-
dimensions—integer· optionalValid values: greater than or equal to 1
Number of dimensions in the embedding output. Configure this only for models that support overriding the output dimensions.
-
ssl_verify—boolean· optional · default:trueIf true, verify the embedding provider's TLS certificate.
-
timeout—integer· optional · default:5000Valid values: greater than or equal to 1
Timeout in milliseconds for requests to the embedding provider.
-
-
azure_openai—object· optionalAzure OpenAI embedding provider settings. Requires
endpointandapi_key.-
endpoint—string· requiredAzure OpenAI embeddings endpoint.
-
api_key—string· requiredAPI key used to authenticate with Azure OpenAI. The value is encrypted with AES before being stored in etcd.
-
dimensions—integer· optionalValid values: greater than or equal to 1
Number of dimensions in the embedding output. Configure this only for models that support overriding the output dimensions.
-
ssl_verify—boolean· optional · default:trueIf true, verify the Azure OpenAI endpoint's TLS certificate.
-
timeout—integer· optional · default:5000Valid values: greater than or equal to 1
Timeout in milliseconds for requests to Azure OpenAI.
-
-
-
vector_search—object· requiredVector search backend configuration.
-
redis—object· requiredRediSearch vector index settings.
-
index—string· optional · default:ai-cacheName of the RediSearch index used by semantic caching.
-
-
-
-
redis_host—string· requiredValid values: at least 2 characters
Address of the Redis server. Required when
policyisredis. -
redis_port—integer· optional · default:6379Valid values: greater than or equal to 1
Port of the Redis server.
-
redis_username—string· optionalUsername for Redis authentication when using Redis ACLs.
-
redis_password—string· optionalPassword for Redis authentication.
In API7 Gateway 3.10.2 or later in the 3.10 release series, and 3.9.16 or later in the 3.9 release series, the value is encrypted with AES256 before being saved to the database.
In APISIX 3.18.0 or later, the value is encrypted with AES before being stored in etcd.
-
redis_database—integer· optional · default:0Valid values: greater than or equal to 0
Database number of the Redis server.
-
redis_timeout—integer· optional · default:1000Valid values: greater than or equal to 1
Timeout in milliseconds for Redis operations.
-
redis_ssl—boolean· optional · default:falseIf true, use TLS for the connection to Redis.
-
redis_ssl_verify—boolean· optional · default:falseIf true, verify the TLS certificate of the Redis server.
-
redis_keepalive_timeout—integer· optional · default:10000Valid values: greater than or equal to 1000
Keepalive timeout in milliseconds for the Redis connection pool. Available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line, and in APISIX from version 3.18.0.
-
redis_keepalive_pool—integer· optional · default:100Valid values: greater than or equal to 1
Keepalive pool size for Redis connections. Available in API7 Enterprise from version 3.9.17 on the 3.9 line and from version 3.10.4 on the 3.10 line, and in APISIX from version 3.18.0.