Last verified 7 min read Inference infrastructure and porting costs

HeyGen's 1.86x on TPU: What the Baseline Actually Is

Google Developers Blog published HeyGen x Google Cloud on 13 August 2026. What the 1.86x is normalised to, and the conditions on the up-to-25% cost line.

This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.

Bottom line: the 1.86x is not a GPU comparison, and the 25% is a ceiling

On 13 August 2026, Google Developers Blog published a joint HeyGen and Google Cloud article, “HeyGen x Google Cloud: Bringing Avatar IV to TPUs”. It describes moving Avatar IV, a video generation stack of more than 18B parameters, onto an eight-chip Trillium (v6e) host, reporting a 1.86x speedup and up to 25% better cost efficiency.

The easy mistake is to assume both figures share a baseline. They do not.

FigureBaseline it is measured againstMetric
1.86xHeyGen’s own first working TPU version (= 1.00x)Relative time per generated video chunk
Up to 25%HeyGen’s own 8xH100 production setupCost efficiency per minute of generated video

The first is not a GPU comparison. The caption of Figure 1 states the baseline outright:

Figure 1. Relative time per generated video chunk, normalized to our first working TPU version (= 1.00×).

Every claim here rests on that single joint article. It is written in HeyGen’s first person and signed by five HeyGen and seven Google Cloud authors, which makes the numbers self-reported benchmarks from the two parties involved. As of 16 August 2026 no independent reporting and no third-party replication could be found.

By condition: the published figures and what they apply to

What the 1.86x covers

PointWhat the article says
DenominatorTheir own first working TPU version
CompositionSharding strategy was already fixed; everything after was kernel and compiler work
MetricTime per chunk, expressed as a relative value
Absolute valuesNot published; no wall-clock figures appear
Quality conditionSame model and same quality gates throughout

The article is explicit about the composition:

The sharding strategy was finalized during the first working version. Everything that followed was kernel and compiler work, and we tracked time per chunk as each change landed.

So the 1.86x is not the effect of moving from GPU to TPU. It is the accumulated effect of tuning after the port already ran.

The sentence the 25% sits in

The result is a pipeline that streams at performance comparable to what we see from our 8×H100 production setup while being up to 25% more cost efficient per minute of generated video.

Three qualifiers cannot be dropped without changing the claim.

QualifierWhat it means
up toA maximum, not a constant 25%
comparablePerformance is on par; the article does not say TPU is faster
BasisUnit prices, discounts, region, utilisation and measurement window are all undisclosed

Hardware and workload conditions

ItemValueSource
Chips8The article
HBM per chip32 GBGoogle Cloud docs (updated 2026-08-11)
bf16 weights, two transformersOver 36 GBThe article
Output resolution720p / 1080pThe article
Frame rate25 fpsThe article
ParametersOver 18BThe article

Avatar IV is not one model. Per the article, three models take turns on every chunk: a diffusion transformer that renders motion conditioned on the audio, a second transformer that super-resolves it, and a VAE decoder that turns latents into pixels.

Steps: quoting these figures without distorting them

  1. Put the baseline on the same line as the number. “1.86x versus their own first working TPU version” is the accurate form
  2. Keep “up to” and “comparable” attached to the 25%. Shortening it to “25% cheaper than 8xH100” leaves the source behind
  3. Attribute it to HeyGen and Google Cloud. With no independent verification, “proven across the industry” is not available
  4. Check whether your workload has the same deadline. The optimisations were chosen against chunk-by-chunk streaming
  5. Do not fill in missing numbers. Porting duration, absolute latency, and TPU or H100 unit prices appear nowhere in the article

Caveats

“86%” is not a speedup

On block alignment for sparse attention, the article says: “The alignment round lifted the kernel from about half of the ceiling this attention shape can reach on the hardware to nearly three-quarters of it, and the rebuilt body closed to about 86%.” Those figures describe how close the kernel got to the ceiling that attention shape can reach on the hardware, not how much faster it became. The speed figure the article does give is “more than ten percent off the super-resolution stage”.

98-99% is measured on their production data

For the optimisation that derives a logit upper bound via the Cauchy-Schwarz inequality and removes the need for an online maximum in softmax, the article reports: “On our production data, 98–99% of heads qualify.” Eligibility is checked per attention head, and heads whose bound is too loose fall back to the standard online path inside the same kernel, because a bound far above the true maximum pushes the exponentials toward underflow. The article does not claim that the same rate holds for other models or workloads.

Generally applicable versus deadline-driven

The split below is this article’s framing, not a classification the source makes. The source does state: “Every optimization here was made against that deadline.” Avatar IV renders and streams chunk by chunk, and if a chunk is late, the video stalls.

MeasureCharacter
Porting through torchax, with production model code running unmodified on JAX arrays and compiled by XLAWorth considering regardless of workload
FSDP weight sharding forced by memory arithmetic, combined with Ulysses sequence parallelism on the same meshSame (follows from over 36 GB against 32 GB per chip)
A two-tier output gate whose first tier requires the delivered video to hash equal to baseline, frame for frameSame
Moving collectives off the critical pathWorks because of the deadline
Short-circuiting softmax with an upper boundWorks because of the deadline
Block alignment for sparse attention (multiples of 128 down to multiples of 16)Works because of the deadline

On collectives, the article notes: “The wire time didn’t shrink. It moved off the critical path, which is all the deadline cares about.” Transfer time itself did not fall. On alignment, roughly one in five live blocks used to come out partial; once blocks were aligned to frame boundaries, the mask predicates, the second pass and the padding all disappeared.

Do not stop at “it ran unmodified”

torchax is a PyTorch frontend on JAX, and the article says production model code ran unmodified through it. The same paragraph also records that TPU-specific work was still required, with each attention variant dispatched to shape-specific Pallas kernels. On rewriting natively in JAX, the article concludes: “The answer was roughly nothing: XLA compiles the whole pipeline end to end either way”.

This is not a full migration announcement

The article refers to “our 8×H100 production setup” in the present tense. Scope, proportion and timing of any migration are not stated.

For related material on how measurement conditions change what a number means, see what GitHub’s official Copilot ROI panel measures and how test-time compute budgets move agent evaluations.

Figures and quotations were checked against the primary sources on 16 August 2026. The source article is 3 days old, the Google Cloud documentation was updated on 11 August 2026, and the HeyGen developer documentation and the torchax repository carry no retrievable publication date, so they reflect the state on 16 August 2026. No retraction or correction was found on that date. Check the official pages for updates.

Sources

  1. HeyGen x Google Cloud: Bringing Avatar IV to TPUs (Google Developers Blog) developers.googleblog.com published 2026-08-13 accessed 2026-08-16
  2. Google Developers Blog index (publication date check) developers.googleblog.com published 2026-08-13 accessed 2026-08-16
  3. TPU v6e (Trillium) - Google Cloud TPU documentation docs.cloud.google.com published 2026-08-11 accessed 2026-08-16
  4. Avatar IV - HeyGen Documentation developers.heygen.com accessed 2026-08-16
  5. google/torchax (official repository) github.com accessed 2026-08-16