Commit Graph

765 Commits

Author SHA1 Message Date
Miguel Ángel 30d6f43bdb chore: release v0.8.31 (#3747)
* chore: release v0.8.31

* docs(release): describe the range fix on its own terms
2026-09-07 12:12:35 -04:00
Miguel Ángel 3874990449 chore: release v0.8.30 (#3733) 2026-09-05 23:14:12 -04:00
James Russo be86a1ec7d fix(render): serve engine and producer files through checked descriptors (#3725)
* fix(engine): read served files through checked descriptors

* fix(producer): retain checked files through streamed responses
2026-09-05 21:38:42 -04:00
James Russo 51a3c7e245 fix(engine): isolate WAV staging in a private directory (#3709)
* fix(engine): create WAV staging files exclusively

* fix(engine): isolate WAV staging in a private directory
2026-09-05 13:44:44 -04:00
James Russo e5d1d9bf0e fix(engine): avoid transform regex backtracking (#3704) 2026-09-05 11:53:43 -04:00
James Russo f13037ecd4 fix(engine): isolate chunked encode temporary files (#3680) 2026-09-05 11:15:46 -04:00
miga-heygen ae3d80c30f chore: release v0.8.29 (#3690) 2026-09-04 21:50:46 -04:00
Miguel Ángel 64ce9fdf1f chore: release v0.8.28 (#3689) 2026-09-04 21:01:08 -04:00
James Russo 19dee4cede fix(render): serve media assets with registered content types (#3677)
Port the MIME mappings reported by fix2015 in #1836 to both render file servers.

Co-authored-by: vitalii.semianchuk <fix20152@gmail.com>
2026-09-04 19:54:29 -04:00
Miguel Ángel c0303fc523 fix(engine): fail closed on blocked SRI scripts (#3664) 2026-09-04 23:15:41 +00:00
Miguel Ángel f334b0e735 fix(engine): retry capture screenshot timeouts (#3641) 2026-09-04 23:15:14 +00:00
Miguel Ángel 19ab83f929 chore: release v0.8.27 (#3608) 2026-09-03 00:27:31 -04:00
Miguel Ángel 84ed587f33 chore: release v0.8.26 (#3597) 2026-09-02 10:36:01 -04:00
Miguel Ángel 6b360f56f7 chore: release v0.8.25 (#3595) 2026-09-02 01:29:06 -04:00
James Russo aceaaebd68 chore: release v0.8.24 (#3593) 2026-09-01 21:54:42 -04:00
James Russo c0887b650d fix(producer): transport safe extraction failure metadata (#3592)
* fix(producer): transport safe extraction failure metadata

* refactor(producer): generalize public error metadata

* test(producer): use vendor-neutral media hosts
2026-09-01 21:35:40 -04:00
James Russo db92f8ab59 test(engine): give the 150-track audio mix test a Windows-sized budget (#3587)
The ENAMETOOLONG regression writes 150 real clip files and the mixer
existence-checks each one: ~58ms on Linux, but past vitest's 5s default on
the Windows lane. packages/engine sets no global testTimeout, so heavy
tests here carry an explicit one.

On timeout its abandoned async work kept calling the shared runFfmpegMock
after afterEach cleared it, so the next test saw 5 calls instead of 3 and
lost its queued once-implementations to the leak. mockReset stops an
aborted test from handing leftovers to the next one.
2026-09-01 14:01:00 -04:00
Miguel Ángel 6cbe3fbe90 chore: release v0.8.23 (#3586) 2026-09-01 13:58:14 -04:00
Miguel Ángel 38e356fba4 chore: release v0.8.22 (#3575)
* chore: release v0.8.22

* docs: include encoder retry in v0.8.22 notes

---------

Co-authored-by: James <james.russo@heygen.com>
2026-08-31 22:54:13 -04:00
James Russo 0f7eebd7e4 fix(encoder): signal host interruptions for retry (#3578)
* fix(encoder): signal host interruptions for retry

* fix(encoder): cover all render interruption paths

* fix(encoder): classify HDR pre-extraction drains
2026-08-31 22:07:43 -04:00
Miguel Ángel f3099dcb27 chore: release v0.8.21 (#3570) 2026-08-31 15:16:22 -04:00
Miguel Ángel 724796e2f0 chore: release v0.8.20 (#3555) 2026-08-30 00:31:08 -04:00
Miguel Ángel 0fd70b1d21 chore: release v0.8.19 (#3551) 2026-08-29 13:58:33 -04:00
Miguel Ángel 5cc2f1bef5 chore: release v0.8.18 2026-08-29 15:38:26 +00:00
Miguel Ángel f6de05efec chore: release v0.8.17 2026-08-28 00:35:58 +00:00
miga-heygen e69be30e98 fix(engine): fail render on sub-composition script failures (#3352) (#3528)
When a composition script throws during execution, the GSAP timeline
registration never arrives and pollSubCompositionTimelines times out.
Previously the render continued with a degenerate 2-frame output and
reported success — now it fails loudly.

Two changes:
1. Detect composition script runtime errors in the browser console
   handler and feed them into scriptLoadFailures, triggering the
   existing fail-fast path (same as script load 404s).
2. Make sub_timeline_script_failure a fatal warning in
   applyRenderWarningPolicy, alongside audio_processing_failed.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-08-28 00:21:34 +00:00
miga-heygen 05275c1e8c fix(producer): assert render artifact duration and frame count before commit (#3506)
* fix(producer): assert render artifact duration and frame count before commit

Refuse to publish an artifact that is significantly shorter or has fewer frames

than the capture pipeline just reported. Adds a duration/frame-count gate on top

of the existing readable-non-empty check inside ArtifactTransaction.validate(),

keyed off the values the orchestrator already carries. Closes #3395.

* fix(producer): wire ffprobe frame count into the artifact duration probe

The frame-count gate added in #3395 accepts an expectedFrames value from
the orchestrator, but defaultArtifactDurationProbe was still returning
only durationSeconds - so the wire was half-built and the assertion
short-circuited on undefined for every real render. Forward meta.frames
from ffprobe so the field-packet case the issue names (container duration
correct, stream shorter) is actually caught by the frame-count check,
not just the duration one.

extractMediaMetadata now populates a new frames field from the video
stream's nb_frames tag, returning undefined when the demuxer did not
report one (fragmented MP4, malformed streams, muxes that require
-count_packets). Callers that gate on the count must treat undefined as
no answer; the assertion already does.

The previous CI run (#32589981916) cancelled shard-6 at the 1h job
timeout after bun install failed to extract the aws-cdk-lib tarball
mid-Docker-build - a cache flake, not a code regression. Pushing a
follow-up commit retriggers CI against the now-populated cache layer;
the regression should clear without further code changes.

---------

Co-authored-by: Santhi Prakash <b.santhiprakash@gmail.com>
2026-08-28 00:19:27 +00:00
Miguel Ángel 720ff5ac9c chore: release v0.8.16 2026-08-27 01:32:37 +00:00
miga-heygen ee64c3b116 fix(engine): stop destroying the AAC priming edit list when muxing (#3505)
`muxVideoWithAudio` passed `-avoid_negative_ts make_zero` unless the caller
set `preserveAudioPrimingEditList`. In practice the dominant path is an AAC
sidecar copied into mp4, where that flag is actively harmful: ffmpeg's
default is `auto`, which the mp4/mov muxers (AVFMT_TS_NEGATIVE) already
resolve to `disabled`. Forcing `make_zero` overrides the correct default,
discards the priming edit list the sidecar encode created, shifts the video
start_time forward by one AAC frame and writes an empty video edit at t=0 —
which edit-list-honoring players (QuickTime/Safari) render as a black first
frame.

Verified with ffprobe on a copy mux of a 30fps h264 mp4 and an AAC sidecar:

  with `make_zero`   video start_time 0.066000, elst: [media time -1,
                     dur 5940] + [media time 6000, dur 180000]
                     audio start_time 0.042993, elst: [media time -1, ...]
  without (this fix) video start_time 0.000000, elst: [media time 6000,
                     dur 180000]
                     audio start_time 0.000000, elst: [media time 1024, ...]

The empty leading edit and the offset both disappear, and the audio keeps
its 1024-sample priming edit.

The flag is now never passed for a mux, in any mode. `preserveAudioPrimingEditList`
is part of the exported engine API, so it stays on `MuxVideoWithAudioOptions`
as `@deprecated` and no-op rather than being removed; the two internal callers
that set it (`assembleStage`, distributed `assemble`) drop it.

`buildEncoderArgs` and `streamingEncoder` still pass the flag for video-only
output and are deliberately left alone — those chunks are consumed as
intermediates, not as a delivered mp4/mov.

Fixes #3487

Co-authored-by: Alexandru Mincu <alex@mountsoftware.ro>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 20:38:11 +00:00
Val 97991bbd35 fix(engine,producer,lint): resolve <source> children for media extract and localize (#3238)
Parent src-only scans skipped multi-format <video>/<audio> markup, so those
elements were never extracted, downloaded, or mixed and rendered blank/silent.
Lint now accepts a child <source src> as a resolvable media src.
2026-08-26 20:17:12 +00:00
Miguel Ángel c9f43ebcfb fix(engine): preserve source frame identity above 99,999 (#3503)
* fix(engine): preserve extracted frame identity

* fix(producer): order legacy distributed frames numerically
2026-08-26 12:30:49 -04:00
Rajan Pantha dae5b7b90b fix(engine): treat a sentineled cache entry with no frames as a miss (#3434)
lookupCacheEntry reported a hit purely on the presence of the
.hf-complete sentinel. The sentinel records that extraction finished,
not that the frames survived, so any per-file cleanup that empties the
directory leaves an entry that rehydrates with zero frames.

rehydrateCacheEntry then returns totalFrames: 0, the clip reaches the
coverage gate with nothing, and the render aborts with a message about
capture coverage. Because the poison is on disk rather than in the
composition, every later render of the project fails the same way with
nothing the user can change to fix it.

An entry now counts as a hit only when it carries the sentinel AND
still holds at least one frame file, so an emptied entry re-extracts.
The check is format-agnostic: a hit must be usable whatever extension
the frames carry.

Addresses the cache half of #3372.
2026-08-26 05:36:32 +00:00
Val f52ec1c25f fix(producer,core): honor relative data-start id-refs in render media scheduling (#3252)
compileTimingAttrs/injectDurations used parseFloat, so data-start="intro"
wrote a NaN data-end and extract preferred that over duration; parseNumeric
now skips the id-ref (parseVideoElements already resolves it).

collectRenderMedia's resolveHostWindow likewise read host data-start with
parseFloat, so chained sub-composition slots (data-start="hook") stacked at
0-2s and every scene after the first rendered black. It now resolves host
starts through the shared resolveReferencedStart, matching the media parsers.

Fixes #3361.
2026-08-26 05:36:24 +00:00
miga-heygen a7e8674758 fix(producer): fall back to screenshot capture on drawElement canvas-not-initialized (#3480)
* fix(producer): fall back to screenshot capture on drawElement canvas-not-initialized

The fast-capture drawElement path only special-cased the "No cached
paint record" error to trigger a per-frame screenshot fallback; every
other error (including "drawElement canvas not initialized", seen at
frame 0 on some macOS/Chrome combinations) was rethrown, hard-failing
the whole render even though the docs promise automatic fallback on
incompatible compositions.

Extend the existing fallback branch (in both captureFrameCore and
captureFrameToBufferPipelined) to also catch canvas-not-initialized
errors via a shared isRecoverableDrawElementError predicate, with a
diagnostic message identifying which case triggered the fallback.

Closes #3423

Co-Authored-By: Miga <noreply@anthropic.com>

* fix(producer): address review — tighten error matching, audit batch path, add fallback-ratio guard

* fix(engine): add prepareFrameForCapture to batch screenshot fallback loop

* fix(engine): split canvas-not-initialized from composition-root-missing errors

drawElementService threw the same HF_DE_CANVAS_NOT_INITIALIZED error for
both !canvas and !root. Missing composition root (navigated/broken page)
was classified recoverable and fell back to pageScreenshotCapture, which
captured blank or wrong content silently.

Now:
- !root → HF_DE_COMPOSITION_ROOT_MISSING (not recoverable, hard fail)
- !canvas → HF_DE_CANVAS_NOT_INITIALIZED (recoverable, screenshot fallback)

Split applied at all 3 emit sites (serial, pipelined, batch).

Co-Authored-By: miga-heygen <miguel.sierra_miga@heygen.com>

---------

Co-authored-by: Miguel Ángel <miguel.sierra@heygen.com>
Co-authored-by: Miga <noreply@anthropic.com>
2026-08-26 03:31:01 +00:00
Miguel Ángel 740f7ead89 chore: release v0.8.15 2026-08-26 03:23:41 +00:00
miga-heygen 3202f3fb87 fix(producer): trip DE parallel-router circuit breaker on stalls and hangs (#3479)
Root cause: the per-worker capture calls in captureFrameRange
(parallelCoordinator.ts) take no abort signal of their own, and only
checked `signal.aborted` BEFORE starting each frame — a no-op once a
worker is already awaiting an in-flight call. On WSL2, the native
drawElement/BeginFrame capture call can hang indefinitely at frame 0
with no error. The DE parallel-router's existing stall watchdog
(captureStreamingStage.ts) correctly fires `stallController.abort()`
after HF_DE_STALL_MS, but that abort had no way to reach a
worker already wedged inside a hung capture call — so
executeParallelCapture's Promise.all waited forever, the render hung
indefinitely, and the CLI's circuit breaker (which only runs after
executeRenderJob settles) never got a chance to trip.

Fix: race each per-frame capture call against the signal actually
firing (raceAgainstAbort), the same "can't cancel, only race" pattern
already used by the sequential capture path. Once the watchdog's abort
is observed, the wedged worker rejects, executeParallelCapture settles,
and the existing pinned-fallback retry / "reverted" outcome / circuit
breaker machinery (already correct) runs end to end.

Also widen the CLI breaker's trip condition from the literal string
"reverted" to "not a clean routed success", so any future non-success
outcome the observability layer records also latches the breaker
instead of silently falling through.

Closes #3441

Co-authored-by: Miga <noreply@anthropic.com>
2026-08-25 23:54:49 +00:00
Miguel Ángel 81069fe47f chore: release v0.8.14 (#3474) 2026-08-24 20:03:19 -04:00
Vance Ingalls 3ed971d018 chore: release v0.8.13 2026-08-24 12:51:02 -07:00
Vance Ingalls 2ca578f945 chore: release v0.8.12 (#3457) 2026-08-23 19:54:55 -07:00
Vance Ingalls 1aec3b4a09 fix(engine): harden grouped audio rendering (#3446)
* fix(core): harden audio FX and group identity

* fix(core): address audio group review feedback

* fix(core): align preview transport with grouped audio

* test(core): pin audio group gain ceiling

* fix(core): preserve solo bridge through stack

* fix(engine): harden grouped audio rendering

* docs(engine): explain grouped mix fallback invariant

* test(engine): allow grouped mixes to finish on Windows
2026-08-23 18:09:21 -07:00
Miguel Ángel 32d58a73e3 chore: release v0.8.11 (#3440) 2026-08-23 14:49:13 -04:00
Miguel Ángel 65b2299db2 fix(engine): preserve static dedup across caption runs (#3438)
* fix(engine): preserve authored clip boundaries after normalization

* perf(engine): bound static verification work across caption runs

* fix(core): preserve explicit nonpositive timeline windows
2026-08-23 13:18:18 -04:00
Miguel Ángel 59a69a145b chore: release v0.8.10 (#3426) 2026-08-22 11:16:32 -04:00
Vance Ingalls f6e8e8ddfd chore: release v0.8.9 (#3422) 2026-08-22 05:57:57 -07:00
Vance Ingalls 073b098e21 feat(engine): preflight psnr filter availability and force-fallback to screenshot on missing (#3418)
## What

Adds a one-shot ffmpeg-psnr filter probe at drawElement session bootstrap.
When the resident ffmpeg is missing or lacks libpostproc (no `psnr` filter),
the capture-session router now force-fallbacks to the screenshot capture
path and emits a `de_gate_reason = "ffmpeg_no_psnr_filter"` telemetry
signal via the existing `render_complete` breakdown.

Also tightens `psnrForDiskSample`'s catch: infrastructure-class ffmpeg
failures (ENOENT, "No such filter") no longer silently skip the sample —
they abort the render so the safety net cannot fail-open post-preflight.

## Why

The drawElement self-verify safety net (parallelCoordinator's
`psnrForDiskSample` → `psnrDb`) shells to `ffmpeg -lavfi psnr`. If ffmpeg
is missing, or was compiled without libpostproc (so the `psnr` filter is
absent), every per-sample compare throws. The existing catch swallows the
error and returns `null` — callers treat that as "skip this sample" and
the render completes with the safety net inoperative.

Field signal ( 9/10 CLI feedback, Slack ts=1787380767.210079,
hyperframes 0.8.7, darwin/arm64, tid=93ff9910-2207-45c2-bc1f-54c0b347d4fe):

> "host ffmpeg lacked psnr filter used by drawElement self-verification,
> but render completed."

The user's frames happened to be byte-identical so no visual damage
shipped — but the safety net silently wasn't running. Any future
compositor-damage bug on that host would have shipped straight through.

## How

Two-part fix, both in `packages/engine`:

1. New `utils/psnrFilterAvailability.ts` — cached probe that runs
   `ffmpeg -hide_banner -filters` once per process and word-boundary-
   matches `psnr` in the output. Any failure (ENOENT, non-zero exit,
   timeout, unparseable output) returns `false`; never rejects.

2. Wired into `services/frameCapture.ts` `initDrawElementOrTransparentBackground`
   right after the Chrome capability probe: when useDrawElement resolves
   true and the preflight returns false, set
   `session.deGateReason = "ffmpeg_no_psnr_filter"` (same low-cardinality
   bucket every other DE gate uses; flows through `getCapturePerfSummary`
   → `render_complete.de_gate_reason` in PostHog), emit a stderr warning
   naming what's missing, and call `routeToFallback()` — the same
   fail-graceful shape as the SwiftShader / CSS-effect / at-risk-timeline
   gates. Skipped under `HF_FORCE_DRAWELEMENT=1` (matches the diagnostic
   knob's policy of bypassing every other gate).

Belt-and-braces: `psnrForDiskSample` now discriminates infrastructure-
class failures (ENOENT / "No such filter" / "Unknown filter") from
per-sample noise (readFile races, transient EPERM). Only the former
re-throw — per-sample noise still returns `null` (skipped sample). The
preflight normally catches this at bootstrap; the re-throw covers
ffmpeg-swapped-mid-render.

## Test plan

- [x] Unit tests added:
  `packages/engine/src/utils/psnrFilterAvailability.test.ts` — mocked
  `execFile` covers: `psnr` present → true; `psnr` absent → false; ENOENT
  → false; non-zero exit → false; result memoized + reset works;
  substring-not-word-boundary → false.
- [x] Unit tests added:
  `isFfmpegInfrastructureFailure` in
  `packages/engine/src/services/parallelCoordinator.test.ts` covers
  ENOENT, "No such filter", "Unknown filter", per-sample EACCES, parse
  errors, null/non-object.
- [x] `bun run test` — `packages/engine/src/utils/psnrFilterAvailability.test.ts`
  (6 tests) + `packages/engine/src/services/parallelCoordinator.test.ts`
  (50 tests) + `frameCapture.test.ts` (26 tests) all pass. Pre-existing
  ffprobe test failures (4) on the base commit are unrelated (missing PNG
  fixture bytes — the file is 129 B on disk, likely LFS-stored).
- [x] `bunx tsc --noEmit -p packages/engine/tsconfig.json` — clean.
- [x] `bunx oxlint <files>` — 0 warnings, 0 errors.
- [x] `bunx oxfmt --check <files>` — clean.

Not covered here: an integration test that boots
`initDrawElementOrTransparentBackground` end-to-end. That path is
Puppeteer-driven and has no unit-scale bootstrap harness in the
repository — the pure preflight + pure discriminator coverage above are
what this PR can prove at the vitest layer.

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 04:23:34 -07:00
Vance Ingalls 6f82acf50c chore: release v0.8.8 (#3411) 2026-08-21 19:04:29 -07:00
Miguel Ángel 92a6076807 test(engine): budget the ffmpeg audio-level tests, and make a stall say why (#3410)
`places a delayed track on its authored start` timed out on the windows runner
and failed an unrelated PR, the second time this week an ffmpeg audio test has
done that.

The previous fix raised the budget in `audioMixer.grouping.test.ts`, which was
the file the symptom named. It was the wrong scope: that file was the only audio
suite with explicit timeouts at all. `audioMixer.level.test.ts` had none, so its
two real-ffmpeg tests ran on vitest's 5s default. The failing one takes ~137ms
locally, so the runner is not 36x slower — but 5s was never a budget anyone
chose for a full mix.

Applied to the suite rather than to each test, so there is one home for it, and
scoped to the ffmpeg-gated describe: the sibling parsing suites are pure and
should keep failing fast at 5s.

Headroom alone would only have delayed an undiagnosable failure. The ffmpeg
process timeout is 5 minutes by default, far above any test budget, so a stalled
mix could only ever surface as a bare "Test timed out" with no stderr and no
failing stage. Tests now cap it at 20s and assert through a helper that reports
`failures` instead of collapsing to `expected false to be true`.

Both claims verified rather than asserted: a deliberately 6s test now passes
where the 5s default would have killed it, and forcing the process timeout to
1ms reports `stage: "prepare", reason: "ffmpeg_timeout"` instead of a timeout.

Reviewing with whitespace ignored is much smaller: adding the third argument to
`describe` reindents the suite body, so 131/101 is really 34/4.
2026-08-21 21:31:58 -04:00
Miguel Ángel 9c73e64a07 test(engine): give the audio grouping mixes room, and make a stall say why (#3408)
`a group FX chain fully cutting its members leaves an ungrouped track
untouched` timed out on the windows runner, failing an unrelated PR. The whole
file runs in ~4s locally and that test in ~1.1s, so 30s was not generous — but
the runner is roughly 10x slower and this test drives more ffmpeg than any of
its siblings, two full mixes plus a group FX chain. 30s was the tightest budget
in the package; 60s is what the rest of the ffmpeg-driven engine tests use.

Headroom alone would only have moved the same undiagnosable failure later,
because nothing here could report why. The production ffmpeg process timeout is
5 minutes, far above any test budget, so a stalled mix could only ever surface
as "Test timed out in 30000ms" with no stderr and no failing stage. Tests now
cap it at 20s, and the mix wrapper throws the recorded failures instead of
returning `success: false` into an `expect(...).toBe(true)` that reports
`expected false to be true` and discards the reason.

Verified by forcing the process timeout to 1ms: the failure goes from a 30s
wall-clock timeout to a 150ms error naming the stage, reason and element
(`stage: "prepare", reason: "ffmpeg_timeout", elementId: "a"`).

This does not explain the Windows stall itself, which I could not reproduce on
macOS. It makes the next occurrence report what it was doing.
2026-08-21 19:46:06 -04:00
Miguel Ángel 41af866bcb chore: release v0.8.7 (#3402) 2026-08-21 15:21:20 -04:00
Miguel Ángel 77566a198b test(engine): give the ffmpeg-bound grouping mixes their 30s timeout (#3398)
audioMixer.grouping.test.ts spawns real ffmpeg per assertion and ran on
vitest's 5s default; on slow Windows runners the FX-chain and envelope
cases land right at the line and fail runs that touch nothing in the
engine. The other ffmpeg-bound engine suites (videoFrameExtractor)
already carry a per-test 30_000 timeout; this brings the grouping suite
in line.
2026-08-21 15:01:44 -04:00