Estimating Hy4 Preview cost requires more than multiplying a context-window size by an input rate. A real agent may resend conversation history, ingest tool results, launch parallel branches, and retry after validation or tool failures. A useful budget therefore totals uncached input, cache-hit input, and output for every model call, adds non-model operating costs, and divides by tasks that actually finish.
To test the estimate with your own workload, visit Tencent Cloud TokenHub for current Hy4 Preview access information. Recheck the official international model pricing page before publication or purchasing.
Lock the price, region, and date first
As of August 28, 2026, Tencent Cloud International lists these Hy4 Preview rates for the Singapore region: US$0.834 per 1 million uncached input tokens, US$2.501 per 1 million output tokens, and US$0.042 per 1 million cache-hit input tokens. Every example below uses that dated regional baseline. It is not a promise that another region or a future billing period will use the same rates.
The official model list identifies the model as hy4-preview and lists a 1M-token context window, a 960k maximum input, and a 64k maximum output. A 1M context window is a capacity limit, not a recommended request size and not a flat fee. Cost estimation should use measured calls rather than the maximum specification.
A cost-per-task formula
For model call i, let Uᵢ be input tokens actually charged as uncached input, Hᵢ be input tokens actually charged as cache hits, and Oᵢ be output tokens. Model cost is:
Model cost = Σ[(Uᵢ × 0.834 + Hᵢ × 0.042 + Oᵢ × 2.501) / 1,000,000]
The same input token cannot belong to both Uᵢ and Hᵢ. More importantly, a planning spreadsheet must not assume a cache hit. Apply the cache-hit rate only to usage that the service actually recognizes as cached. This guide makes no assumption about cache retention, matching behavior, or future availability; validate those conditions against current documentation and observed usage.
A broader business formula is:
Cost per completed task = (model + tools + infrastructure + human review) / completed tasks
Search services, browser automation, code sandboxes, database queries, storage, and reviewer time are not included in the three Hy4 token rates. Track them separately. Use completed tasks in the denominator when measuring delivered outcomes, because dividing only by attempts hides the cost of failures.
Example 1: One long-context request
Assumption: A document-analysis call has 300,000 uncached input tokens, 5,000 output tokens, and no cache hit.
- Input:
300,000 × 0.834 / 1M = $0.250200 - Output:
5,000 × 2.501 / 1M = $0.012505 - Call total:
$0.250200 + $0.012505 = $0.262705
Output is smaller in token count here, but its unit rate is higher. Repetitive summaries, verbose intermediate narration, or unnecessary reasoning text can therefore remain a meaningful cost driver.
Example 2: A measured partial cache hit
Assumption: One task makes two calls. Call one has 200,000 uncached input tokens and 4,000 output tokens. Call two has 20,000 new uncached input tokens, 180,000 cache-hit input tokens, and 4,000 output tokens.
- Uncached input:
220,000 × 0.834 / 1M = $0.183480 - Cache-hit input:
180,000 × 0.042 / 1M = $0.007560 - Output:
8,000 × 2.501 / 1M = $0.020008 - Task total:
$0.211048
For comparison, if both 200,000-token inputs were entirely uncached, the calculation would be 400,000 × 0.834 / 1M + $0.020008 = $0.353608. The conditional difference is $0.142560. It is not a universal savings percentage: changed prefixes, misses, and a different request mix produce different results.
Example 3: Multi-turn context accumulation
Assumption: A three-turn task has actual input counts of 40,000, 55,000, and 70,000 tokens. Outputs are 2,000, 2,000, and 3,000 tokens. All input is uncached. Later calls carry part of the prior conversation.
- Total input:
165,000 × 0.834 / 1M = $0.137610 - Total output:
7,000 × 2.501 / 1M = $0.017507 - Model total:
$0.155117
Using only the last turn's 70,000-token input would omit 95,000 tokens processed earlier. Multiplying the maximum window by three would overstate usage. Aggregate the measured input for each call.
Example 4: Tool results and a failed retry
When logs, web results, or test output are inserted into a later model request, they become part of that request's input. Any fee charged by the tool provider remains a separate line item.
Assumption: An agent's three normal steps use 30,000, 45,000, and 60,000 input tokens, with 2,000 output tokens per step. A tool then fails. A fourth model call includes the error and uses 62,000 input tokens plus 2,000 output tokens. All input is uncached.
- Three normal steps:
135,000 × 0.834 / 1M + 6,000 × 2.501 / 1M = $0.127596 - Retry increment:
62,000 × 0.834 / 1M + 2,000 × 2.501 / 1M = $0.056710 - Total with retry:
$0.184306
Discarding the failed result does not make its model usage free. Bound maximum steps, retries by error class, total tokens, elapsed time, and spend. Also distinguish failure rate from retry rate: some failures stop immediately, while others trigger one or several new calls.
Example 5: Parallel branches and synthesis
Assumption: A research agent launches two branches. Branch A uses 50,000 input and 2,000 output tokens; branch B uses 80,000 and 3,000; a synthesis call uses 25,000 and 2,000. Everything is uncached.
- Input:
155,000 × 0.834 / 1M = $0.129270 - Output:
7,000 × 2.501 / 1M = $0.017507 - Total:
$0.146777
Parallelism can reduce wall-clock duration, but billing still reflects all calls that ran. Branch B belongs in the numerator even if the synthesizer rejects its conclusion. Concurrency itself is not a token-price multiplier; extra calls, duplicated context, and branch-specific retries are the cost sources.
Example 6: Human review and completed-task economics
Assumption: A monthly workflow attempts 10,000 tasks. Model cost per attempt matches Example 4 at $0.184306, and an external tool costs $0.006 per attempt. Eight percent of attempts need four minutes of review at an assumed internal rate of $36 per hour. The workflow completes 9,600 tasks.
- Model:
10,000 × $0.184306 = $1,843.06 - External tool:
10,000 × $0.006 = $60.00 - Human review:
800 × 4 / 60 × $36 = $1,920.00 - Total:
$3,823.06 - Cost per attempt:
$3,823.06 / 10,000 = $0.382306 - Cost per completed task:
$3,823.06 / 9,600 ≈ $0.398235
The tool rate, review share, reviewer rate, and completion rate are explicit assumptions—not Tencent Cloud prices or claimed customer results. The example shows why inference per request and business cost per result are different metrics. Review can exceed model expense, while unsuccessful attempts reduce the number of delivered outcomes.
Instrument the workload without inventing billing fields
Do not design around a field name that has not been verified in the current API or billing documentation. Instead, require an auditable application record that can be reconciled with the official bill: task identifier, call order, model and region, uncached input, cache-hit input, output, branch, failure reason, retry relationship, external tool expense, human minutes, and completion status. The exact source and name for each value should be confirmed during implementation.
Segment the results by workload rather than relying on one average. Repository analysis, long-document review, and short question answering have different tail behavior. Track cost per attempt, cost per completion, observed cache-hit share, incremental retry cost, human-handoff rate, and P95 task cost. Run sensitivity ranges for input growth, output caps, hit rate, failure rate, and branch count, but keep forecasts separate from actuals.
Build a three-case proof-of-concept budget
A procurement estimate should not present one highly precise number without its assumptions. Create low, baseline, and high cases for the same task. The low case can use measured shorter inputs, controlled outputs, and only cache-hit volume already demonstrated in tests. The baseline can use sample medians and an observed retry mix. The high case should include P95 input and output, additional branches, poor cache reuse, and more human escalation. Keep the formula constant, change the inputs explicitly, and place the pricing date and region at the top of the table.
Suppose the baseline task begins with the $0.184306 model cost from Example 4, and measurements show that 12% of tasks incur one additional retry costing $0.056710. Under the explicit constraint of at most one such retry, estimated average model cost is $0.184306 + 12% × $0.056710 = $0.1911112. If failures can trigger two or more retries, model each retry-count cohort separately. Applying 12% once would understate the tail.
Set a proof-of-concept circuit breaker for spend, steps, and elapsed time. When a task reaches a limit, stop launching branches and transfer the trace for human review. The objective is not to erase failed work. Its incurred model and tool expense stays in the numerator, while analysis determines whether the trigger was context growth, cache misses, an unstable tool, or an overly strict acceptance check.
Three distinctions keep the dashboard honest. A cache-hit rate does not mean all input receives the $0.042 rate; new and missed input remains uncached. Parallel execution is not automatically more expensive if it replaces a longer sequential search, but comparison must use total calls. Human review does not belong on the Hy4 token bill, yet it still belongs in business cost per completed task.
Hy4 Preview's long context, caching, and Function Calling can support demanding agent workloads, but they do not automatically make those workloads economical. A defensible estimate sums every measured call, preserves failed and discarded work in the numerator, separates non-model costs, and normalizes by completed tasks. Because the model remains in Preview, recheck pricing, specifications, and operational behavior before each production or procurement decision.