High Performance
Application Load Balancer (ALB) adopts a fully managed application-layer cluster architecture. A single instance can handle millions of queries per second (QPS) and tens of millions of concurrent connections, making it capable of supporting high-concurrency business scenarios such as e-commerce platforms, online education, and interactive entertainment with daily visits exceeding 10 million. ALB supports HTTP/2 multiplexing and the QUIC low-latency protocol, providing a stable entry point for your high-traffic business.
High Availability
ALB adopts a multi-availability zone (AZ) deployment mode by default, where AZs serve as hot backups for each other. If any AZ fails, traffic is automatically switched within seconds, achieving an availability of 99.995%. Combined with Auto Scaling (AS), ALB automatically distributes traffic to healthy instances when an instance in a real server group becomes abnormal, helping ensure business continuity. The cross-AZ disaster recovery capability enables your core services to obtain enterprise-grade high availability without additional investment. Fine-Grained Scheduling
ALB implements fine-grained traffic forwarding based on application-layer protocol characteristics. It supports routing matching based on multi-dimensional conditions such as the domain name, URL path, HTTP Header, Query String, and Cookie. This meets the requirements of complex scenarios in microservices architectures, including canary releases, A/B testing, and multi-tenant isolation. ALB also provides the request redirection and URL rewriting features. You can adjust business logic without the need to modify backend code, significantly reducing iteration costs.
Rich Protocol Support
ALB provides comprehensive coverage of common communication protocols for modern applications, supporting HTTP, HTTPS, HTTP/2, WebSocket, and QUIC, with native support for the gRPC framework to seamlessly meet high-performance communication requirements between microservices. The QUIC protocol is particularly suitable for latency-sensitive business scenarios such as Tencent Real‑Time Communication (TRTC) and online gaming. The server-sent events (SSE) streaming capability can support real-time inference result pushing for large language mode (LLM)‑powered AI applications. A single load balancing solution can support diverse workloads including web applications, mobile APIs, and AI inference.
Security and Reliability
ALB integrates a multi-layered security protection system, supports TLS 1.3 encrypted transmission across the entire link, and has built-in SNI multi-domain certificate management capabilities. A single listener can be associated with multiple certificates to meet large-scale HTTPS business requirements. You can use Web Application Firewall (WAF) and Anti-DDoS Pro to help identify and defend against web attacks such as SQL injection and cross‑site scripting (XSS) attacks, as well as DDoS threats, building a defense-in-depth system from the network layer to the application layer for backend services. Cloud-Native Integration
ALB can be used as a cloud-native Ingress gateway and is deeply integrated with Tencent Kubernetes Engine (TKE) and Serverless Cloud Function (SCF). It supports declarative management of routing rules via Ingress resources, providing unified traffic entry control for containerized applications. When your TKE dynamically scales in or out via AS, ALB automatically detects backend node changes and updates forwarding policies in real time. You can achieve elastic traffic scheduling in cloud-native environments without manual intervention. Elastic Billing
ALB adopts a capacity unit (CU) billing mode and charges on a pay-as-you-go basis based on actual resource consumption, including new connections, concurrent connections, processed traffic, and rule evaluations. Compared with fixed-specification instances, this mode aligns more closely with the actual usage costs of elastic workloads. With pay-as-you-go billing based on actual resource consumption, costs decrease during business troughs, and capacity automatically scales out during business peaks without the need to pre-estimate capacity. This helps you effectively optimize infrastructure costs while maintaining performance.