HeyGen's 1.86x on TPU: What the Baseline Actually Is
Google Developers Blog published HeyGen x Google Cloud on 13 August 2026. What the 1.86x is normalised to, and the conditions on the up-to-25% cost line.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Bottom line: the 1.86x is not a GPU comparison, and the 25% is a ceiling
On 13 August 2026, Google Developers Blog published a joint HeyGen and Google Cloud article, “HeyGen x Google Cloud: Bringing Avatar IV to TPUs”. It describes moving Avatar IV, a video generation stack of more than 18B parameters, onto an eight-chip Trillium (v6e) host, reporting a 1.86x speedup and up to 25% better cost efficiency.
The easy mistake is to assume both figures share a baseline. They do not.
| Figure | Baseline it is measured against | Metric |
|---|---|---|
| 1.86x | HeyGen’s own first working TPU version (= 1.00x) | Relative time per generated video chunk |
| Up to 25% | HeyGen’s own 8xH100 production setup | Cost efficiency per minute of generated video |
The first is not a GPU comparison. The caption of Figure 1 states the baseline outright:
Figure 1. Relative time per generated video chunk, normalized to our first working TPU version (= 1.00×).
Every claim here rests on that single joint article. It is written in HeyGen’s first person and signed by five HeyGen and seven Google Cloud authors, which makes the numbers self-reported benchmarks from the two parties involved. As of 16 August 2026 no independent reporting and no third-party replication could be found.
By condition: the published figures and what they apply to
What the 1.86x covers
| Point | What the article says |
|---|---|
| Denominator | Their own first working TPU version |
| Composition | Sharding strategy was already fixed; everything after was kernel and compiler work |
| Metric | Time per chunk, expressed as a relative value |
| Absolute values | Not published; no wall-clock figures appear |
| Quality condition | Same model and same quality gates throughout |
The article is explicit about the composition:
The sharding strategy was finalized during the first working version. Everything that followed was kernel and compiler work, and we tracked time per chunk as each change landed.
So the 1.86x is not the effect of moving from GPU to TPU. It is the accumulated effect of tuning after the port already ran.
The sentence the 25% sits in
The result is a pipeline that streams at performance comparable to what we see from our 8×H100 production setup while being up to 25% more cost efficient per minute of generated video.
Three qualifiers cannot be dropped without changing the claim.
| Qualifier | What it means |
|---|---|
| up to | A maximum, not a constant 25% |
| comparable | Performance is on par; the article does not say TPU is faster |
| Basis | Unit prices, discounts, region, utilisation and measurement window are all undisclosed |
Hardware and workload conditions
| Item | Value | Source |
|---|---|---|
| Chips | 8 | The article |
| HBM per chip | 32 GB | Google Cloud docs (updated 2026-08-11) |
| bf16 weights, two transformers | Over 36 GB | The article |
| Output resolution | 720p / 1080p | The article |
| Frame rate | 25 fps | The article |
| Parameters | Over 18B | The article |
Avatar IV is not one model. Per the article, three models take turns on every chunk: a diffusion transformer that renders motion conditioned on the audio, a second transformer that super-resolves it, and a VAE decoder that turns latents into pixels.
Steps: quoting these figures without distorting them
- Put the baseline on the same line as the number. “1.86x versus their own first working TPU version” is the accurate form
- Keep “up to” and “comparable” attached to the 25%. Shortening it to “25% cheaper than 8xH100” leaves the source behind
- Attribute it to HeyGen and Google Cloud. With no independent verification, “proven across the industry” is not available
- Check whether your workload has the same deadline. The optimisations were chosen against chunk-by-chunk streaming
- Do not fill in missing numbers. Porting duration, absolute latency, and TPU or H100 unit prices appear nowhere in the article
Caveats
“86%” is not a speedup
On block alignment for sparse attention, the article says: “The alignment round lifted the kernel from about half of the ceiling this attention shape can reach on the hardware to nearly three-quarters of it, and the rebuilt body closed to about 86%.” Those figures describe how close the kernel got to the ceiling that attention shape can reach on the hardware, not how much faster it became. The speed figure the article does give is “more than ten percent off the super-resolution stage”.
98-99% is measured on their production data
For the optimisation that derives a logit upper bound via the Cauchy-Schwarz inequality and removes the need for an online maximum in softmax, the article reports: “On our production data, 98–99% of heads qualify.” Eligibility is checked per attention head, and heads whose bound is too loose fall back to the standard online path inside the same kernel, because a bound far above the true maximum pushes the exponentials toward underflow. The article does not claim that the same rate holds for other models or workloads.
Generally applicable versus deadline-driven
The split below is this article’s framing, not a classification the source makes. The source does state: “Every optimization here was made against that deadline.” Avatar IV renders and streams chunk by chunk, and if a chunk is late, the video stalls.
| Measure | Character |
|---|---|
| Porting through torchax, with production model code running unmodified on JAX arrays and compiled by XLA | Worth considering regardless of workload |
| FSDP weight sharding forced by memory arithmetic, combined with Ulysses sequence parallelism on the same mesh | Same (follows from over 36 GB against 32 GB per chip) |
| A two-tier output gate whose first tier requires the delivered video to hash equal to baseline, frame for frame | Same |
| Moving collectives off the critical path | Works because of the deadline |
| Short-circuiting softmax with an upper bound | Works because of the deadline |
| Block alignment for sparse attention (multiples of 128 down to multiples of 16) | Works because of the deadline |
On collectives, the article notes: “The wire time didn’t shrink. It moved off the critical path, which is all the deadline cares about.” Transfer time itself did not fall. On alignment, roughly one in five live blocks used to come out partial; once blocks were aligned to frame boundaries, the mask predicates, the second pass and the padding all disappeared.
Do not stop at “it ran unmodified”
torchax is a PyTorch frontend on JAX, and the article says production model code ran unmodified through it. The same paragraph also records that TPU-specific work was still required, with each attention variant dispatched to shape-specific Pallas kernels. On rewriting natively in JAX, the article concludes: “The answer was roughly nothing: XLA compiles the whole pipeline end to end either way”.
This is not a full migration announcement
The article refers to “our 8×H100 production setup” in the present tense. Scope, proportion and timing of any migration are not stated.
For related material on how measurement conditions change what a number means, see what GitHub’s official Copilot ROI panel measures and how test-time compute budgets move agent evaluations.
Figures and quotations were checked against the primary sources on 16 August 2026. The source article is 3 days old, the Google Cloud documentation was updated on 11 August 2026, and the HeyGen developer documentation and the torchax repository carry no retrievable publication date, so they reflect the state on 16 August 2026. No retraction or correction was found on that date. Check the official pages for updates.
Sources
この記事の日本語版: HeyGen's 1.86x on TPU: What the Baseline Actually Is(日本語)