Tab | Description | Applicable Scenarios |
Global Overview | Displays gateway-level aggregated metrics and Top lists | Quickly understand the overall operational status of the gateway |
LLM Monitoring | Displays specialized metrics for large model calls. | Analyze the performance and consumption of LLM requests. |
MCP Monitoring | Displays specialized metrics for MCP tool calls. | Monitor the health status of MCP Server and Tool. |
Metric Card | Description |
Total number of requests | Total number of gateway requests within the selected time range |
Average request latency P95 | The 95th percentile of end-to-end request latency (ms) |
Maximum request latency P99 | The 99th percentile of end-to-end request latency (ms) |
Error rate (%) | Proportion of failed requests to total requests (HTTP ≥ 400) |
Token consumption | Total consumption of input + output tokens |
First Token latency TTFT P95 | The 95th percentile of the time to first Token for streaming responses (ms) |
Panel Name | Description |
Request volume trend (QPS) | Shows the trend of requests per second. |
Error rate trend (%) | Shows the trend of error rate over time. |
End-to-end latency distribution P50/P95/P99 | Shows the distribution trend of request latency percentiles (ms). |
First Token latency TTFT P50/P95/P99 | Shows the percentile trend of first Token latency in streaming responses (ms). |
Token consumption trend | Shows the stacked trend chart of input/output tokens. |
Panel Name | Description | Purpose |
Top Token consumption (by model API) | Top model APIs by Token consumption | Identify high-consumption APIs and optimize costs. |
Top request volume (by consumer) | Top consumers by request volume | Understand traffic distribution and allocate quotas appropriately. |
Top average request latency (by model API) | Top model APIs by average latency | Locate high-latency APIs and optimize performance. |
Top error rate (by model API) | Top model APIs by error rate | Quickly identify abnormal APIs and troubleshoot issues. |
Top First Token latency TTFT (by model API) | Top model APIs by TTFT latency | Optimize streaming response experience. |
Metric Card | Description |
Total number of requests | Total number of LLM requests |
P95 latency | The 95th percentile of end-to-end latency (ms) |
P99 latency | The 99th percentile of end-to-end latency (ms) |
Error rate (%) | Error rate of LLM requests |
First Token latency P95 | The 95th percentile of TTFT for streaming responses (ms) |
Token consumption | Total Token consumption of LLM requests |
Total Input Token | Total input tokens |
Total Output Token | Total output tokens |
Token generation rate (tok/s) | Average number of tokens generated per second |
Input Cache hit rate (%) | Proportion of tokens hit by Prompt Cache |
Number of LLM HTTP requests | Total number of LLM HTTP requests |
Number of Chat operations | Number of Chat Completion requests |
Panel Name | Description |
Latency percentile trend (P50/P95/P99) | Shows the trend of LLM request latency percentiles. |
Input/Output Token trend (stacked) | Shows the stacked trend of input/output tokens. |
Prompt length distribution (Input Token histogram) | Shows the distribution of Prompt lengths. |
TTFT First Token latency P50/P95/P99 | Shows the percentile trend of first Token latency in streaming responses. |
TTFT distribution histogram | Shows the distribution histogram of TTFT. |
TTFT P95 grouped by Provider | Shows TTFT P95 grouped by different providers. |
Provider response time P95 | Backend response time P95 for each provider |
Provider throughput (number of requests) | Request throughput per provider |
Top 10 call traffic | Top 10 consumers/APIs by call volume |
Input Cache Hit Rate reflects the hit status of the model-side Prompt Cache. A higher hit rate indicates a greater number of duplicate Prompts. You can further reduce costs by enabling the AI Gateway cache.Metric Card | Description |
Tool invocation count | Total number of MCP Tool invocations |
Tool success rate | Success rate of Tool invocations |
Tool P99 latency | The 99th percentile of Tool execution duration (ms) |
Panel Name | Description |
Tool invocation trend (grouped by tool) | Trend of invocation counts per Tool |
Success rate/parameter error rate trend (%) | Trend of Tool invocation success rate and parameter error rate |
P99 latency trend by tool | Trend of P99 latency per Tool |
Average latency trend | Trend of average latency for Tool invocations |
Tool invocation detail ranking (by tool) | Ranking of number of invocations, success rate, and latency details per Tool |
Was this page helpful?
You can also Contact sales or Submit a Ticket for help.
Help us improve! Rate your documentation experience in 5 mins.
Feedback