tencent cloud

Cloud Native Intelligent Gateway

Log Dashboard

Download
Focus Mode
Font Size
Last updated: 2026-09-22 18:51:11
AI-Translated

Scenarios

The AI Gateway integrates with Tencent Cloud Log Service (CLS) to provide LLM and MCP monitoring and visualization capabilities based on log dashboards. Through the CLS log dashboard, you can use the following capabilities:
Global Overview: View gateway-level core metrics, including total number of requests, latency percentiles, error rate, Token consumption, and first Token latency.
LLM Monitoring: Perform in-depth analysis of performance metrics for large model invocations, including Token consumption trends, Provider response times, and Prompt length distribution.
MCP Monitoring: Monitors the success rate, latency, and invocation trends of MCP tool calls.
Note:
The data in the log dashboard originates from AI Gateway access logs that have been delivered to CLS. Before using it, ensure that you have enabled the delivery of LLM access logs and MCP access logs in Data Observation > Log Delivery.

Prerequisites

1. The AI gateway instance has been created and is in a running state.
2. You have configured and enabled CLS log shipping in Data Observation > Log Delivery (the log topic has been created and data is being written normally).
3. The gateway already has LLM or MCP request traffic (data generation begins approximately 1-2 minutes after delivery is enabled).

Viewing the Log Dashboard

1. Log in to the Microservices Platform console. In the left sidebar, select AI Gateway to go to the instance list.
2. In the instance list, click the name of the target instance to go to its details page.
3. In the left sidebar, select Data Observation and then click the Log Dashboard Tab.
The Log Dashboard page contains three sub-tabs:
Tab
Description
Applicable Scenarios
Global Overview
Displays gateway-level aggregated metrics and Top lists
Quickly understand the overall operational status of the gateway
LLM Monitoring
Displays specialized metrics for large model calls.
Analyze the performance and consumption of LLM requests.
MCP Monitoring
Displays specialized metrics for MCP tool calls.
Monitor the health status of MCP Server and Tool.

Use Cases

Scenario 1: Global Overview

The Global Overview provides an aggregation of gateway-level core metrics, helping you quickly grasp the overall operational status.
Note:
Currently, the Global Overview only includes metrics related to LLM calls and does not yet include metrics for MCP and Agent calls.

1. Core Metric Cards

Metric Card
Description
Total number of requests
Total number of gateway requests within the selected time range
Average request latency P95
The 95th percentile of end-to-end request latency (ms)
Maximum request latency P99
The 99th percentile of end-to-end request latency (ms)
Error rate (%)
Proportion of failed requests to total requests (HTTP ≥ 400)
Token consumption
Total consumption of input + output tokens
First Token latency TTFT P95
The 95th percentile of the time to first Token for streaming responses (ms)

2. Viewing Quick Exception Discovery

The Rapid Anomaly Detection (Consumer Risk Top 10) panel displays the ranking of high-risk consumers:
The abnormal score from the consumer risk calculation is comprehensively evaluated based on dimensions such as error rate, latency, and Token consumption.
This helps you quickly locate problematic consumers and promptly investigate anomalies.

3. Viewing the Distribution and Trend Panel

Panel Name
Description
Request volume trend (QPS)
Shows the trend of requests per second.
Error rate trend (%)
Shows the trend of error rate over time.
End-to-end latency distribution P50/P95/P99
Shows the distribution trend of request latency percentiles (ms).
First Token latency TTFT P50/P95/P99
Shows the percentile trend of first Token latency in streaming responses (ms).
Token consumption trend
Shows the stacked trend chart of input/output tokens.

4. Viewing Top Lists

Panel Name
Description
Purpose
Top Token consumption (by model API)
Top model APIs by Token consumption
Identify high-consumption APIs and optimize costs.
Top request volume (by consumer)
Top consumers by request volume
Understand traffic distribution and allocate quotas appropriately.
Top average request latency (by model API)
Top model APIs by average latency
Locate high-latency APIs and optimize performance.
Top error rate (by model API)
Top model APIs by error rate
Quickly identify abnormal APIs and troubleshoot issues.
Top First Token latency TTFT (by model API)
Top model APIs by TTFT latency
Optimize streaming response experience.

Scenario 2: Viewing LLM Monitoring

LLM Monitoring provides specialized metrics for large model invocations, helping you perform in-depth analysis of the performance characteristics of LLM requests.

1. Core Metric Cards

Metric Card
Description
Total number of requests
Total number of LLM requests
P95 latency
The 95th percentile of end-to-end latency (ms)
P99 latency
The 99th percentile of end-to-end latency (ms)
Error rate (%)
Error rate of LLM requests
First Token latency P95
The 95th percentile of TTFT for streaming responses (ms)
Token consumption
Total Token consumption of LLM requests
Total Input Token
Total input tokens
Total Output Token
Total output tokens
Token generation rate (tok/s)
Average number of tokens generated per second
Input Cache hit rate (%)
Proportion of tokens hit by Prompt Cache
Number of LLM HTTP requests
Total number of LLM HTTP requests
Number of Chat operations
Number of Chat Completion requests

2. Viewing the LLM Trend and Distribution Panel

Panel Name
Description
Latency percentile trend (P50/P95/P99)
Shows the trend of LLM request latency percentiles.
Input/Output Token trend (stacked)
Shows the stacked trend of input/output tokens.
Prompt length distribution (Input Token histogram)
Shows the distribution of Prompt lengths.
TTFT First Token latency P50/P95/P99
Shows the percentile trend of first Token latency in streaming responses.
TTFT distribution histogram
Shows the distribution histogram of TTFT.
TTFT P95 grouped by Provider
Shows TTFT P95 grouped by different providers.
Provider response time P95
Backend response time P95 for each provider
Provider throughput (number of requests)
Request throughput per provider
Top 10 call traffic
Top 10 consumers/APIs by call volume
Note:
The Input Cache Hit Rate reflects the hit status of the model-side Prompt Cache. A higher hit rate indicates a greater number of duplicate Prompts. You can further reduce costs by enabling the AI Gateway cache.

Scenario 3: Viewing MCP Monitoring

MCP Monitoring provides specialized metrics for MCP Tool calls, helping you monitor the health status of MCP Servers and Tools.

1. Core Metric Cards

Three core metric cards are displayed at the top of the MCP Monitoring page.
Metric Card
Description
Tool invocation count
Total number of MCP Tool invocations
Tool success rate
Success rate of Tool invocations
Tool P99 latency
The 99th percentile of Tool execution duration (ms)

2. Viewing the MCP Trend and Ranking Panel

Panel Name
Description
Tool invocation trend (grouped by tool)
Trend of invocation counts per Tool
Success rate/parameter error rate trend (%)
Trend of Tool invocation success rate and parameter error rate
P99 latency trend by tool
Trend of P99 latency per Tool
Average latency trend
Trend of average latency for Tool invocations
Tool invocation detail ranking (by tool)
Ranking of number of invocations, success rate, and latency details per Tool

Scenario 4: Custom Analysis in the CLS Console

When the preset panels on the Log Dashboard cannot meet your analysis requirements, you can navigate to the CLS console to perform custom analysis.

Operation Steps

1. In the upper-right corner of the Log Dashboard page, click View More in CLS.
2. The system will automatically navigate to the corresponding log topic in the Tencent Cloud CLS console.
3. In the CLS console, you can:
Use SQL statements to perform custom queries on log data.
Create custom dashboards and charts.
Configure log alarm rules.
Export log data for offline analysis.


Help and Support

Was this page helpful?

Help us improve! Rate your documentation experience in 5 mins.

Feedback