Files
vercel__workflow/scripts/aggregate-worlds-data.mjs
Pranay Prakash 5dd15452cd Test and benchmark community worlds against e2e tests (#482)
* Add new github workflow

* Enable pull_request trigger for community worlds workflow

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Add community worlds manifest and generation scripts

- Add community-worlds.json manifest as single source of truth
- Add scripts/generate-community-worlds-workflow.mjs to generate CI workflow
- Add scripts/generate-community-worlds-docs.mjs to generate docs section
- Update aggregate-benchmarks.js to load community worlds dynamically
- Add pnpm generate:community-worlds script
- Update docs/deploying/world/index.mdx with community worlds

The manifest-based approach allows:
- E2E tests to be auto-generated from the manifest
- Benchmark aggregation to include community worlds
- Docs to stay in sync with tested worlds

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Fix YAML syntax error - quote strings starting with @

The @ symbol has special meaning in YAML, so package names like
@workflow-worlds/turso need to be quoted.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Fix Redis health-cmd quoting, remove unpublished starter world

- Quote health-cmd when it contains spaces (fixes Docker arg parsing)
- Remove @workflow-worlds/starter as it's not published to npm

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Add benchmarks and summary job for community worlds

- Add build job to share artifacts between benchmark jobs
- Add benchmark jobs for Turso, MongoDB, and Redis worlds
- Update summary job to show both E2E and benchmark status matrix
- Add left border/indent to sidebar child items for visual hierarchy
- Update workflow generator to support benchmark generation

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Reuse build artifacts for E2E tests

E2E jobs now depend on the shared build job and download
artifacts instead of rebuilding packages from scratch.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Add Worlds Ecosystem dashboard to docs

- Create worlds-manifest.json with official and community worlds
- Add aggregate-worlds-data.mjs script for processing E2E and benchmark results
- Create WorldsDashboard, WorldCard, and BenchmarkChart components
- Add /docs/worlds page showing compatibility status and performance
- Include sample data for development

The dashboard shows:
- E2E test progress per world (pass/fail/skip counts)
- Benchmark performance comparison across all worlds
- Filter by official vs community worlds

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Fix CI to output JSON test results and add Jazz world

- Update workflow generator to output JSON test results from vitest
- Upload E2E results as artifacts for parsing in summary job
- Summary job now shows actual pass/fail/skip counts per world
- Add Jazz world to worlds-manifest.json (requires external credentials)
- Add update-worlds-status.yml workflow to auto-update dashboard data
- Update TypeScript types to support null lastRun and metrics

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Refactor community worlds to use reusable workflows

Instead of a generated workflow file, integrate community world testing
directly into tests.yml and benchmarks.yml using reusable workflows.

- Add reusable workflows for E2E tests: e2e-community-world.yml (no services),
  e2e-community-world-mongodb.yml, e2e-community-world-redis.yml
- Add reusable workflows for benchmarks: benchmark-community-world.yml,
  benchmark-community-world-mongodb.yml, benchmark-community-world-redis.yml
- Update tests.yml to call reusable workflows for Turso, MongoDB, Redis
- Update benchmarks.yml to include community world benchmarks in summary
- Delete generated community-worlds.yml and generator script

This approach:
- Inherits proper Rust/SWC setup from the main workflows
- Keeps all CI in the established patterns
- Makes adding new community worlds straightforward

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Fix benchmark timing file naming for community worlds

Add WORKFLOW_BENCH_BACKEND env var support to bench.bench.ts so community
world benchmarks generate timing files with the correct backend suffix
(e.g., bench-timings-nextjs-turbopack-turso.json instead of -local.json).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Add @workflow-worlds/starter to community worlds test matrix

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Unify worlds manifest and add dynamic GitHub API fetching

- Merge community-worlds.json into worlds-manifest.json with type field
- Add server-side data fetching from GitHub API for worlds dashboard
- Remove static worlds-status.json, fetch CI artifacts dynamically
- Update all references to use unified manifest format
- Remove obsolete community-worlds.yml and update-worlds-status.yml workflows

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Skip community worlds for non-nextjs-turbopack in benchmark summary

Community worlds only run against nextjs-turbopack, so hide the
"missing" rows for Express and Nitro frameworks in the benchmark
comparison tables.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Fix stream benchmark detection to check for actual TTFB data

The previous check `!== null` incorrectly returned true for undefined,
causing all benchmarks to show TTFB columns. Now explicitly checks
for a number type.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Hide Worlds Ecosystem page from sidebar

The page is still accessible via direct link at /docs/worlds but
won't appear in the navigation until it's been further iterated on.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Fix review comments: trailing newline and division by zero

- Add trailing newline when replacing Community Worlds section in docs
- Fix division by zero in WorldCard benchmark calculation when metrics is empty

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Consolidate community world workflows with service-type parameter

- Create setup-workflow-dev composite action for common setup steps
- Add service-type input to benchmark-community-world.yml and e2e-community-world.yml
- Use conditional job execution (if: inputs.service-type == 'mongodb') to handle different services
- Update benchmarks.yml and tests.yml to pass service-type parameter
- Delete redundant workflow files:
  - benchmark-community-world-mongodb.yml
  - benchmark-community-world-redis.yml
  - e2e-community-world-mongodb.yml
  - e2e-community-world-redis.yml

Reduces workflow files from 11 to 7 and eliminates ~500 lines of duplicated YAML.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Generate community world test matrix from worlds-manifest.json

Replace hardcoded community world jobs with dynamic matrix generation using
scripts/create-community-worlds-matrix.mjs. This allows adding/removing
community worlds by editing the manifest instead of multiple workflow files.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Apply suggestion from @vercel[bot]

Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>

* Add Samples column and separate local/production benchmarks

- Add Samples column to all benchmark tables showing iteration count
- Separate benchmark results into Local Development and Production sections
- Add explanatory context for each section (localhost vs Vercel deployment)
- Add GitHub action step summaries to e2e community world tests
- Create aggregate-e2e-results.js script for parsing vitest JSON output

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Remove obsolete generate-community-worlds-docs script

The Worlds Ecosystem page now fetches from worlds-manifest.json at runtime,
making this script unnecessary. The npm script also referenced a non-existent
workflow generator script.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Maximize composite action usage and reorganize benchmark output

- Update setup-workflow-dev composite action with optional Rust, install-dependencies, and install-args inputs
- Update tests.yml to use composite action in unit, e2e-vercel-prod, getTestMatrix, e2e-local-*, and getCommunityWorldsMatrix jobs
- Update benchmarks.yml to use composite action in build, benchmark-local, benchmark-postgres, benchmark-vercel, and getCommunityWorldsMatrix jobs
- Reorganize benchmark output to group by benchmark test with local/production tables within each benchmark
- Remove invalid $schema reference from worlds-manifest.json

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Add beads stealth mode stuff (for personal claude memory - will remove stealth if people want)

* Add E2E test results PR comment summary

- Add pr-comment-start job to create/update PR comment when tests start
- Add artifact uploads to all e2e test jobs (vercel-prod, local-dev, local-prod, local-postgres, windows)
- Update e2e-community-world.yml with consistent artifact naming (e2e-community-*)
- Add summary job to aggregate all e2e results and update PR comment
- Extend aggregate-e2e-results.js with --mode aggregate for multi-job PR summary
- Group results by category (Vercel Production, Local Development, etc.)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Add step summaries to all e2e test jobs

Add "Generate E2E summary" step to each individual e2e job:
- e2e-vercel-prod
- e2e-local-dev
- e2e-local-prod
- e2e-local-postgres
- e2e-windows

Each job now outputs pass/fail/skip counts to GITHUB_STEP_SUMMARY.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Consolidate community world workflows to single job

Replace 3 mutually exclusive jobs (e2e/e2e-mongodb/e2e-redis) with a single
job that starts services via docker run when needed. This eliminates the
skipped job entries that appear in the GitHub Actions UI.

- Use conditional docker run steps instead of services: block
- Add health check loops to wait for service readiness
- Add cleanup step to stop containers

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* Fix extractWorldId to handle community world artifact naming

Add handling for `e2e-results-community-{world}` pattern so community
world test results are properly extracted (e.g., `e2e-results-community-turso`
now correctly extracts `turso` instead of `community-turso`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
2025-12-02 15:12:44 -08:00

354 lines
9.8 KiB
JavaScript

#!/usr/bin/env node
/**
* Aggregates E2E test results and benchmark data from CI runs
* into a unified worlds-status.json for the dashboard.
*
* Usage:
* node scripts/aggregate-worlds-data.mjs [results-dir] [--output path/to/output.json]
*
* Input files expected:
* - e2e-results-{world}.json: Vitest JSON output for E2E tests
* - bench-results-{app}-{world}.json: Vitest benchmark output
* - bench-timings-{app}-{world}.json: Custom timing data
*
* Output:
* - worlds-status.json: Combined status of all worlds
*/
import fs from 'fs';
import path from 'path';
import { fileURLToPath } from 'url';
const __dirname = path.dirname(fileURLToPath(import.meta.url));
const rootDir = path.join(__dirname, '..');
// Parse command line arguments
const args = process.argv.slice(2);
let resultsDir = '.';
let outputPath = path.join(rootDir, 'docs/public/data/worlds-status.json');
for (let i = 0; i < args.length; i++) {
if (args[i] === '--output' && args[i + 1]) {
outputPath = args[i + 1];
i++;
} else if (!args[i].startsWith('--')) {
resultsDir = args[i];
}
}
// Load worlds manifest
const manifestPath = path.join(rootDir, 'worlds-manifest.json');
let manifest;
try {
manifest = JSON.parse(fs.readFileSync(manifestPath, 'utf-8'));
} catch (e) {
console.error(`Error: Could not load worlds manifest: ${e.message}`);
process.exit(1);
}
// Get all worlds from manifest (now a flat array with type field)
function getAllWorlds() {
return manifest.worlds || [];
}
// Find all E2E result files
function findE2EResultFiles(dir) {
const files = [];
try {
const entries = fs.readdirSync(dir, { withFileTypes: true });
for (const entry of entries) {
const fullPath = path.join(dir, entry.name);
if (entry.isDirectory()) {
files.push(...findE2EResultFiles(fullPath));
} else if (
entry.name.startsWith('e2e-results-') &&
entry.name.endsWith('.json')
) {
files.push(fullPath);
}
}
} catch (e) {
// Directory may not exist
}
return files;
}
// Find all benchmark result files
function findBenchmarkFiles(dir) {
const files = [];
try {
const entries = fs.readdirSync(dir, { withFileTypes: true });
for (const entry of entries) {
const fullPath = path.join(dir, entry.name);
if (entry.isDirectory()) {
files.push(...findBenchmarkFiles(fullPath));
} else if (
entry.name.startsWith('bench-timings-') &&
entry.name.endsWith('.json')
) {
files.push(fullPath);
}
}
} catch (e) {
// Directory may not exist
}
return files;
}
// Parse E2E result file (vitest JSON output)
function parseE2EResults(filePath) {
try {
const data = JSON.parse(fs.readFileSync(filePath, 'utf-8'));
let total = 0;
let passed = 0;
let failed = 0;
let skipped = 0;
const tests = [];
// Vitest JSON format
for (const file of data.testResults || []) {
for (const test of file.assertionResults || []) {
total++;
const status = test.status;
if (status === 'passed') passed++;
else if (status === 'failed') failed++;
else if (status === 'skipped' || status === 'pending') skipped++;
tests.push({
name: test.fullName || test.title,
status: status === 'pending' ? 'skipped' : status,
duration: test.duration || 0,
});
}
}
// Alternative format (direct vitest output)
if (total === 0 && data.numTotalTests) {
total = data.numTotalTests;
passed = data.numPassedTests || 0;
failed = data.numFailedTests || 0;
skipped = (data.numPendingTests || 0) + (data.numTodoTests || 0);
}
return { total, passed, failed, skipped, tests };
} catch (e) {
console.error(
`Warning: Could not parse E2E results ${filePath}: ${e.message}`
);
return null;
}
}
// Parse benchmark timing file
function parseBenchmarkTimings(filePath) {
try {
const data = JSON.parse(fs.readFileSync(filePath, 'utf-8'));
const metrics = {};
if (data.summary) {
for (const [benchName, stats] of Object.entries(data.summary)) {
metrics[benchName] = {
mean: stats.avgExecutionTimeMs,
min: stats.minExecutionTimeMs,
max: stats.maxExecutionTimeMs,
samples: stats.samples,
};
// Add TTFB for stream benchmarks
if (stats.avgFirstByteTimeMs !== undefined) {
metrics[benchName].ttfb = {
mean: stats.avgFirstByteTimeMs,
min: stats.minFirstByteTimeMs,
max: stats.maxFirstByteTimeMs,
};
}
}
}
return metrics;
} catch (e) {
console.error(
`Warning: Could not parse benchmark timings ${filePath}: ${e.message}`
);
return null;
}
}
// Extract world ID from filename
function extractWorldFromFilename(filename, prefix) {
// e2e-results-{world}.json -> world
// bench-timings-{app}-{world}.json -> world
const basename = path.basename(filename, '.json');
const withoutPrefix = basename.replace(prefix, '');
// For bench files, format is {app}-{world}, we want the last part
const parts = withoutPrefix.split('-');
return parts[parts.length - 1];
}
// Aggregate all data
function aggregateWorldsData() {
const allWorlds = getAllWorlds();
const timestamp = new Date().toISOString();
// Initialize worlds status
const worldsStatus = {};
for (const world of allWorlds) {
worldsStatus[world.id] = {
type: world.type,
name: world.name,
package: world.package,
description: world.description,
docs: world.docs,
repository: world.repository,
e2e: null,
benchmark: null,
};
}
// Process E2E results
const e2eFiles = findE2EResultFiles(resultsDir);
for (const file of e2eFiles) {
const worldId = extractWorldFromFilename(file, 'e2e-results-');
if (worldsStatus[worldId]) {
const results = parseE2EResults(file);
if (results) {
const progress =
results.total > 0
? Math.round((results.passed / results.total) * 1000) / 10
: 0;
worldsStatus[worldId].e2e = {
status:
results.failed === 0
? 'passing'
: results.passed > 0
? 'partial'
: 'failing',
total: results.total,
passed: results.passed,
failed: results.failed,
skipped: results.skipped,
progress,
tests: results.tests,
lastRun: timestamp,
};
}
}
}
// Process benchmark results
const benchFiles = findBenchmarkFiles(resultsDir);
const benchmarksByWorld = {};
for (const file of benchFiles) {
const worldId = extractWorldFromFilename(file, 'bench-timings-');
if (!benchmarksByWorld[worldId]) {
benchmarksByWorld[worldId] = {};
}
const metrics = parseBenchmarkTimings(file);
if (metrics) {
// Merge metrics (could have multiple apps)
Object.assign(benchmarksByWorld[worldId], metrics);
}
}
// Assign benchmarks to worlds
for (const [worldId, metrics] of Object.entries(benchmarksByWorld)) {
if (worldsStatus[worldId]) {
worldsStatus[worldId].benchmark = {
status: Object.keys(metrics).length > 0 ? 'measured' : 'pending',
metrics,
lastRun: timestamp,
};
}
}
return {
$schema: './worlds-status.schema.json',
lastUpdated: timestamp,
commit: process.env.GITHUB_SHA || null,
branch: process.env.GITHUB_REF_NAME || null,
worlds: worldsStatus,
};
}
// Generate test matrix (detailed per-test breakdown)
function generateTestMatrix(worldsStatus) {
const allTests = new Map(); // testName -> { world -> status }
// Collect all tests from all worlds
for (const [worldId, world] of Object.entries(worldsStatus)) {
if (world.e2e?.tests) {
for (const test of world.e2e.tests) {
if (!allTests.has(test.name)) {
allTests.set(test.name, {});
}
allTests.get(test.name)[worldId] = test.status;
}
}
}
// Convert to array format
const tests = [];
for (const [name, results] of allTests) {
tests.push({ name, results });
}
// Sort by test name
tests.sort((a, b) => a.name.localeCompare(b.name));
return {
lastUpdated: new Date().toISOString(),
tests,
};
}
// Main
console.log('Aggregating worlds data...');
console.log(` Results directory: ${resultsDir}`);
console.log(` Output path: ${outputPath}`);
const status = aggregateWorldsData();
const testMatrix = generateTestMatrix(status.worlds);
// Ensure output directory exists
const outputDir = path.dirname(outputPath);
if (!fs.existsSync(outputDir)) {
fs.mkdirSync(outputDir, { recursive: true });
}
// Write worlds-status.json
fs.writeFileSync(outputPath, JSON.stringify(status, null, 2));
console.log(`\nGenerated ${outputPath}`);
// Write test-matrix.json
const testMatrixPath = path.join(path.dirname(outputPath), 'test-matrix.json');
fs.writeFileSync(testMatrixPath, JSON.stringify(testMatrix, null, 2));
console.log(`Generated ${testMatrixPath}`);
// Summary
const worldCount = Object.keys(status.worlds).length;
const withE2E = Object.values(status.worlds).filter((w) => w.e2e).length;
const withBenchmarks = Object.values(status.worlds).filter(
(w) => w.benchmark
).length;
console.log(`\nSummary:`);
console.log(` Total worlds: ${worldCount}`);
console.log(` With E2E data: ${withE2E}`);
console.log(` With benchmark data: ${withBenchmarks}`);
for (const [id, world] of Object.entries(status.worlds)) {
const e2eStatus = world.e2e
? `${world.e2e.passed}/${world.e2e.total} (${world.e2e.progress}%)`
: 'no data';
const benchStatus = world.benchmark
? `${Object.keys(world.benchmark.metrics).length} benchmarks`
: 'no data';
console.log(` - ${world.name}: E2E ${e2eStatus}, Benchmarks ${benchStatus}`);
}