* fix(producer): mix audio into a container that can record encoder delay Every rendered composition's audio landed 1024 samples (21.33 ms at 48 kHz) after its authored `data-start`, against a frame-accurate video track. The mix is AAC-encoded, and AAC encoders emit ~1024 priming samples. The mix was written to a raw ADTS `.aac` file, which has nowhere to record that delay, so it decoded as real leading silence and every stage downstream preserved it faithfully. Measuring each intermediate localises it precisely: the source WAV is exact, the mixer's own output is already 21.33 ms late, and the pad/trim and mux stages inherit it unchanged. The filter graph itself is correct - run by hand to PCM it lands on the authored start. Switch the artifact to an MP4-family container, which stores the delay as an edit list that decoders strip. Same codec, same bitrate, so no size or quality change. The filename is a contract shared by three consumers - the mux input, the distributed plan artifact, and the PNG-sequence sidecar handed to users for NLE ingest - and its extension is what selects the muxer. Give it one owner in the engine rather than five literals, so those consumers cannot drift onto different containers. Note for reviewers: this renames the distributed plan's audio artifact, which is an on-disk contract between the plan writer and the assembler. Both move together here, but a plan written by an older build would not be found by a newer assembler. Flagging in case that mixed-version window matters for how these are deployed. * fix(cloud): read the plan audio artifact name from the producer contract The aws-lambda and gcp-cloud-run adapters each restated the plan's audio filename in five places, so renaming it in the producer left them looking for a file that is no longer written. CI caught it: the gcp dispatch test asserting a plan has no audio artifact started seeing one. Export the name from `@hyperframes/producer/distributed` and consume it in both adapters. This is the same failure the constant exists to prevent, one package boundary further out: a literal that drifts from the writer's is a silently missing audio track rather than a loud error, because both call sites only ever ask whether the file exists. * fix(cloud): accept a legacy plan's audio artifact name for one release Review raised a rolling-deploy window I had flagged but left undecided: `plan` and `assemble` are separate invocations bridged by object storage, so a pre-rollout planner can be paired with a post-rollout assembler. Both readers locate the artifact by existence alone, which makes that pairing a silently muted video rather than an error. That is reachable enough to be worth two lines, so reads now accept the old name while writes only ever emit the new one. Give the fallback one owner (`resolvePlanAudioPath` / `isPlanAudioArtifactPath`) rather than four call sites, marked for deletion one release out. Also fixes a hole in the first pass of this: the plan-v2 materializer matched either name but then joined the CURRENT one, so a legacy plan resolved to a path that was never written. It now joins the artifact's own name. Review nits in the same pass: correct the pad-branch docstring, which still described a concat-copy shape the pad branch stopped using when it moved to apad + re-encode, and fix the Windows fixture's stale `.aac` output extension so it cannot model a shape that reintroduces the priming delay. * test(producer): rebake the missing-host-comp-id golden without the audio delay The pinned reference was rendered before this branch, so it carries the 1024 sample encoder-priming delay in its audio. With the delay gone the correct audio now sits ahead of the reference and the harness's envelope correlation drops below its floor. Cross-correlating the old and new references at native 48 kHz gives a lag of exactly 1024 samples (21.33 ms) at a correlation of 0.99985: same audio, moved by exactly the amount this branch removes. Regenerated inside the CI container (Dockerfile.test, ffmpeg 5.1.9) rather than natively, so the reference matches the encoder CI will compare against - the container reproduced CI's failure to the digit (correlation 0.3938764027803616, lagWindows -12) before the rebake and passes at correlation 1.0 after it. Note for archaeology: the new reference is also 3 dB louder than the old one. That gap is not from this branch - `main` and this branch render the fixture at the same level - it is pre-existing drift the reference had accumulated, which a scale-invariant correlator could never see. The rebake absorbs it. Only output.mp4 is updated. `--update` also rewrites compiled.html, but that diff is embedded-font churn with no bearing on the comparison, which reports "Failed at compilation: 0" either way. * test(producer): rebake the variables-prod golden without the audio delay Same cause as the missing-host-comp-id rebake, caught by shard-8 once the earlier shard stopped failing and the rest of the matrix could run: this reference also carries the encoder-priming delay this branch removes. Reproduced in the CI container to the digit (correlation 0.42704173048439215, lagWindows -12), rebaked there, and it now passes at correlation 1.0. Worth recording: the shift here is 2048 samples (42.67 ms) at correlation 0.99983, exactly twice the 1024 of the other fixture. The delay compounds once per un-compensated AAC generation, and this fixture's audio needs its duration normalized, so it takes the pad/trim branch's re-encode and picks up a second frame of priming on top of the mixer's. So the pre-fix error was not a fixed 21 ms - it grew with the number of times the audio was re-encoded. All nine shards ran in that CI round with only this one failing, so the matrix has now covered every fixture against this change.
@hyperframes/gcp-cloud-run
Google Cloud Run + Cloud Workflows adapter for HyperFrames distributed
rendering. The OSS render primitives (plan → renderChunk × N →
assemble) are pure functions over local file paths; this package is the
deployment, orchestration, and storage glue that runs them on Google Cloud —
the GCP counterpart to @hyperframes/aws-lambda.
Two surfaces, one package:
- Server-side handler (
./server) — a Cloud Run HTTP service that dispatchesplan/renderChunk/assembleon the request body'sActionfield, bridging GCS ↔ the container's filesystem around each OSS primitive. This is what the bundledDockerfileruns. - Client-side SDK (
./sdk) —renderToCloudRun,getRenderProgress,deploySite,validateDistributedRenderConfig, andcomputeRenderCost. Call these from a Node process (CI, CLI, app backend) to drive a deployed stack without writing GCS / Workflows boilerplate.
The package is not a dependency of @hyperframes/producer; install it
separately.
Architecture
GCS bucket ←→ Cloud Run service (plan / renderChunk / assemble)
▲
│ OIDC-authenticated http.post, one per step
│
Cloud Workflows (Plan → parallel RenderChunk → Assemble)
- Plan downloads the project tarball and publishes either a legacy v1 planDir tarball or a v2 manifest plus content-addressed artifacts.
- RenderChunk runs in a parallel
forloop in the workflow, fanned out up to the plan's chunk count. Each invocation renders one chunk and uploads it. - Assemble downloads every chunk + audio, stitches the final deliverable, and uploads it.
Every step is a POST to the same Cloud Run URL with a different Action.
The workflow accumulates each step's small result body and returns
{ Plan, Chunks, Assemble } so getRenderProgress can read frame totals and
per-step durations on success.
Plan transport selection
Plan v2 is recommended for new integrations. renderToCloudRun still
interprets an omitted planProtocol as "v1" for backwards compatibility,
so new callers should select v2 explicitly:
await renderToCloudRun({
// ...project, bucket, workflow, service, and config...
planProtocol: "v2",
});
V2 uses separate manifest and content-addressed artifact locators throughout the workflow. Unknown protocols and integrity failures fail closed; a render never mixes v1 and v2 artifacts.
Chrome runtime
Unlike the Lambda adapter — which fights a 250 MB ZIP ceiling and
decompresses @sparticuz/chromium into /tmp at runtime — Cloud Run runs a
container image. The Dockerfile installs the same pinned
chrome-headless-shell build and font set the production renderer uses, at a
fixed path, and exports HYPERFRAMES_CHROME_PATH. CDP-level BeginFrame
support is a binary/runtime capability, so the image build launches that
exact executable and requires an enable + warm-up + PNG-returning
HeadlessExperimental.beginFrame probe to pass. The end-to-end smoke also
requires every chunk to report effective CaptureMode: "beginframe", which
catches runtime fallback separately from build-time packaging. There is no
runtime decompression step and no packaging ceiling.
Deploying
The terraform/ module provisions everything: the GCS render bucket, the
Cloud Run service, the Cloud Workflows definition, two least-privilege
service accounts (the service reads/writes the bucket; the workflow invokes
the service), and a runaway-request alert.
# 1. Build + push the image (Cloud Build or local docker).
gcloud builds submit . \
--tag REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG
# 2. Apply the module.
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform init
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform apply \
-var project_id=PROJECT \
-var region=us-central1 \
-var image=REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG
Terraform outputs render_bucket_name, service_url, workflow_name, and
region — pass them straight into the SDK.
Using the SDK
import { renderToCloudRun, getRenderProgress } from "@hyperframes/gcp-cloud-run/sdk";
const handle = await renderToCloudRun({
projectDir: "./my-composition",
config: { fps: 30, width: 1920, height: 1080, format: "mp4" },
bucketName: "hyperframes-render-my-project", // from terraform output
projectId: "my-project",
location: "us-central1",
workflowId: "hyperframes-render",
serviceUrl: "https://hyperframes-render-abc.us-central1.run.app",
});
// Poll until done.
let progress = await getRenderProgress({ executionName: handle.executionName });
while (progress.status === "running") {
await new Promise((r) => setTimeout(r, 5000));
progress = await getRenderProgress({ executionName: handle.executionName });
}
console.log(progress.status, progress.outputFile, progress.costs.displayCost);
deploySite is called implicitly when you pass projectDir; call it
yourself to pre-upload once and reuse the siteHandle across many renders
(e.g. personalised template batches).
Running tests
bun test # unit tests over an in-memory GCS double — no network
bun run typecheck
The live end-to-end smoke (build image → terraform apply → render a fixture
through the workflow → PSNR-compare → destroy) lives at
examples/gcp-cloud-run/scripts/smoke.sh and needs a GCP project with
billing enabled.
What's still ahead
- Mid-flight per-chunk progress.
getRenderProgressreports coarserunningprogress and exact numbers on success. Reading the Cloud Workflows step-entries API would give per-chunk progress while the render is in flight; tracked as a follow-up. - Cloud Run Jobs / Firebase Functions variants. This first version targets Cloud Run services + Workflows (the closest analog to Lambda + Step Functions). The same handler runs unchanged under Cloud Run Jobs; only the orchestration trigger differs.