Metric Name | Metric Meaning |
Total number of requests | Total number of requests, summed based on the selected time granularity |
Average request latency | Average request latency, calculated as the average based on the selected time granularity. |
Maximum request latency | Maximum request latency, calculated as the maximum value based on the selected time granularity. |
Number of requests directly returned by the gateway | Number of requests that are not forwarded to the backend but directly responded to by the gateway (for example, when authentication fails or throttling is triggered), summed based on the selected time granularity. |
Average gateway latency | Average time taken by the gateway itself to process requests. |
Maximum gateway latency | Maximum time taken by the gateway itself to process requests. |
Number of 2xx requests | Number of successful requests (for example, 200 OK) sent from the client to the AI gateway, summed based on the selected time granularity. |
Number of 3xx requests | Number of requests redirected after they are sent from the client to the AI gateway, summed based on the selected time granularity. |
Number of 4xx requests | Number of client errors directly returned by the gateway for illegal requests sent from the client to the AI gateway (for example, due to authentication failure or exceeding the throttling limit), such as 401 authentication failure, 403 insufficient permissions, and 429 throttling, summed based on the selected time granularity. |
Number of 5xx requests | Number of server-side errors returned by the backend service after the AI gateway forwards messages to it (such as 500 backend exception, 502 backend invalid response, and 504 backend unreachable), summed based on the selected time granularity. |
Number of 403 requests | Permission error. The backend is capable of processing the request but denies authorization access. |
Number of 404 requests | Number of requests that failed to reach the backend service because the requested resource was not found on the backend server, summed based on the selected time granularity. |
Number of 429 requests | Number of requests failed to be sent to the backend service because the requests are throttled, summed based on the selected time granularity |
Number of 499 requests | Number of requests that failed to reach the backend service because the client actively disconnected before the backend responded, summed based on the selected time granularity. |
Number of 502 requests | Number of errors where the gateway receives an invalid response from the backend server (usually due to a connection failure) while attempting to execute a backend request, summed based on the selected time granularity. |
Number of 504 requests | Number of errors where the backend machine is unreachable when the gateway attempts to execute a backend request, summed based on the selected time granularity. |
Number of requests forwarded to the backend by the gateway | Number of requests successfully forwarded by the gateway to the backend service, summed based on the selected time granularity. |
Average backend latency | Average time taken by the backend service to process requests, calculated as the average based on the selected time granularity. |
Maximum backend latency | Maximum time taken by the backend service to process requests, calculated as the maximum value based on the selected time granularity. |
Number of backend 2xx requests | Number of successful requests (for example, 200 OK) processed by the backend service, summed based on the selected time granularity. |
Number of backend 3xx requests | Number of redirection requests processed by the backend service, summed based on the selected time granularity. |
Number of backend 4xx requests | Number of illegal requests sent to the backend service, summed based on the selected time granularity. |
Number of backend 5xx requests | Number of server-side errors returned by the backend service (for example, 500 backend exception, 502 backend invalid response, 504 backend unreachable), summed based on the selected time granularity. |
Number of backend 404 requests | Number of errors where the requested backend service resource is not found on the backend server, summed based on the selected time granularity. |
Number of backend 429 requests | Number of errors where the backend service request fails because the request is throttled, summed based on the selected time granularity. |
Number of backend 499 requests | Number of errors where the backend service request fails because the client actively disconnects before the backend responds, summed based on the selected time granularity. |
Number of backend 502 requests | Number of errors where the backend service request fails because the backend service receives an invalid response, summed based on the selected time granularity. |
Number of backend 504 requests | Number of errors where the backend service request fails because the backend machine is unreachable, summed based on the selected time granularity. |
Metric Name | Metric Meaning |
Number of LLM HTTP requests | The number of HTTP calls initiated by the gateway to the large model provider. This metric directly reflects the invocation frequency of the model API. |
Total tokens consumed by LLM | The total number of tokens consumed by the gateway from the large model provider, which is the sum of the actual tokens consumed for input (Prompt) and output (Completion). It is used to evaluate the total data throughput of Token consumption. |
Tokens consumed by LLM prompt | The total number of tokens consumed by the model for the input (Prompt) part when the large model processes a request. |
Tokens consumed by LLM completion | The total number of tokens consumed by the model for the output (Completion) part when the large model generates a response. This metric is one of the core bases for evaluating model invocation costs. |
Average latency per LLM request (ms) | The average duration from when the gateway sends a request to the model provider to when it receives the complete response. This metric reflects the end-to-end response performance of the model provider. |
Average latency per token by LLM provider (ms) | The average time spent by the model provider to consume each Token. This metric reflects the Token consumption speed of the model provider. |
Metric Name | Metric Meaning |
MCP request quantity | Total number of requests received by the MCP service within the selected time range. This metric directly reflects the invocation frequency of the MCP Server. |
Average MCP request latency (ms) | Average time taken by the MCP service to process requests (in milliseconds). This metric reflects the performance of the MCP Server. |
MCP request success rate (%) | Percentage of successful requests among total requests invoked by the MCP service. This metric reflects the stability of the MCP Server. |
Was this page helpful?
You can also Contact sales or Submit a Ticket for help.
Help us improve! Rate your documentation experience in 5 mins.
Feedback