Continuation of #26108. Previously we built the whole output as a single rope, and then flattened it in one go at the end. This is fine for smaller files, but for larger files it has 2 problems: 1. Each `output += segment` append allocates a 32-byte cons string cell. When printing a large AST, so many of these are allocated that they can fill "new space" and trigger garbage collection. When that happens, all these little cons strings have to be copied into the other half of new space, which is slow. Even worse, if they survive 2 rounds of GC, they get promoted into "old space" where (a) they may get scattered across memory, ruining cache coherence, and (b) they sit there, taking up memory, until the next major GC round. 2. Flattening a rope string which contains any segments which are 2-byte strings (non-Latin1 characters) is way more expensive than flattening a rope which is pure Latin1/ASCII. If a single segment contains Unicode, the entire flatten operation deopts and becomes several times slower. This PR attacks both these problems by flattening a large output in chunks. - When output reaches 16 KiB, flatten that as a chunk and store it in an array. - This leaves all the little cons string cells unreferenced, so garbage collection will discard them all, instead of copying them across to keep them alive. - If file contains some non-Latin1 characters, that does still heavily impact the performance of flattening the chunk they're in, but only that chunk - other chunks which are pure ASCII still take the fast path. The 16 KiB limit is a heuristic, chosen after experimentation with different values. ### Benchmarks The effect of this change is minimal on small files, but dramatic on large ones (negative is faster): | Fixture | Bytes | Change | |:---| ---:| ---:| | `tiny.js` | 26 | \+3.2% | | `RadixUIAdoptionSection.jsx` | 2,424 | \-0.1% | | `react.development.js` | 50,496 | \+2.0% | | `binder.ts` | 126,212 | \-1.9% | | `lodash.js` | 182,262 | \-32.5% | | `App.tsx` | 298,130 | \-26.4% | | `kitchen-sink.tsx` | 662,560 | \-14.2% | | `antd.js` | 5,100,647 | \-28.4% | ### Total cost of flattening After this PR, compared to where we were prior to #26106 (positive is slower): | Fixture | Bytes | Change | |:---| ---:| ---:| | `tiny.js` | 26 | \+51.5% | | `RadixUIAdoptionSection.jsx` | 2,424 | \+134.2% | | `react.development.js` | 50,496 | \+35.5% | | `binder.ts` | 126,212 | \+39.0% | | `lodash.js` | 182,262 | \+45.3% | | `App.tsx` | 298,130 | \+55.0% | | `kitchen-sink.tsx` | 662,560 | \+51.3% | | `antd.js` | 5,100,647 | \+2.2% | The effect of flattening is still very negative, but (as mentioned in #26106) it was always a cost, just one that we were ignoring. At least we've now managed to mitigate it. In the case of really large files (`antd.js`), benchmark is pretty much unchanged from before flattening came into the picture. But now our measure _includes_ the large flattening cost which was previously _on top of_ the cost of printing - so the real cost for the user is much reduced from where it was previously. ### Sourcemap builds Printing with sourcemaps on large files actually gets faster after this change. It was already paying the cost of string flattening, now it pays less. Printing `lodash.js` with sourcemaps is about 15% faster now.
5.1 KiB
oxc-codegen
Fast, synchronous code generation for JavaScript and TypeScript ASTs.
oxc-codegen turns an ESTree or
TS-ESTree AST into formatted source
code. It supports JavaScript, JSX, TypeScript, and TSX.
The printer is a port of Oxc's Rust oxc_codegen crate. With the default options, both printers
produce byte-identical output: tab indentation, double-quoted strings, and no comments.
Installation
npm install oxc-codegen
oxc-codegen is ESM-only and requires Node.js ^20.19.0 or >=22.12.0.
Quick start
Pair it with oxc-parser to parse and print source code:
import { printSync } from "oxc-codegen";
import { parseSync } from "oxc-parser";
const { program } = parseSync("input.js", "const answer=6*7");
const { code } = printSync(program);
console.log(code);
// const answer = 6 * 7;
You can also print a manually constructed AST:
const program = {
type: "Program",
sourceType: "script",
body: [
{
type: "ExpressionStatement",
expression: {
type: "CallExpression",
callee: {
type: "MemberExpression",
object: { type: "Identifier", name: "console" },
property: { type: "Identifier", name: "log" },
computed: false,
optional: false,
},
arguments: [{ type: "Literal", value: "Hello!" }],
optional: false,
},
},
],
};
console.log(printSync(program).code);
// console.log("Hello!");
TypeScript and TSX
Set ts when the AST can contain TypeScript nodes. For TSX, set both ts and jsx:
const { program } = parseSync("component.tsx", "const Box = <T,>(value: T) => <div>{value}</div>");
const { code } = printSync(program, {
ts: true,
jsx: true,
});
API
printSync(node, options?)
function printSync(
node: ESTree.Program | ESTree.Statement,
options?: Options,
): {
code: string;
map: SourceMap | null;
};
Prints a complete Program or a single statement and returns the generated source code,
and (when requested) a standard Source Map v3 object.
import { printSync } from "oxc-codegen";
import { parseSync } from "oxc-parser";
const sourceText = "const answer=6*7";
const { program } = parseSync("input.js", sourceText);
const { code, map } = printSync(program, {
sourcemap: true,
sourceFilename: "input.js",
sourceText,
});
Source-map mappings require sourceText and nodes with valid Oxc start / end offsets.
A manually constructed AST without offsets can still be printed, but its source map has
an empty mappings string.
Options
| Option | Type | Default | Description |
|---|---|---|---|
indent |
string |
"\t" |
Non-empty string of spaces and/or tabs used for one indent level |
startingIndentLevel |
number |
0 |
Starting indent level, from 0 to 1000 |
jsx |
boolean |
false |
Enable TSX-safe printing for ambiguous TypeScript syntax |
ts |
boolean |
false |
Select the printer that supports TypeScript nodes |
sourcemap |
boolean |
false |
Return a Source Map v3 object in map |
sourceFilename |
string |
"" |
Original source filename recorded in the source map |
sourceText |
string |
- | Original text required for source-map mappings and content |
Why pure JavaScript?
Most Oxc packages use native bindings. This package deliberately does not: when an AST already
lives in JavaScript, passing the entire object graph across a JS/native boundary can cost more than
printing it in place. oxc-codegen avoids that serialization and uses specialized printer builds
for JavaScript and TypeScript workloads.
See DESIGN.md for the implementation details and performance constraints.
Current limitations
- Comments are not printed.
- Minified output is not supported.
Benchmarks
Representative time per printSync call:
| Fixture | Bytes | Time |
|---|---|---|
tiny.js |
26 | 0.0001 ms |
RadixUIAdoptionSection.jsx |
2,518 | 0.0070 ms |
react.development.js |
72,141 | 0.1518 ms |
binder.ts |
193,077 | 0.3364 ms |
App.tsx |
415,340 | 1.2912 ms |
lodash.js |
544,096 | 0.7977 ms |
kitchen-sink.tsx |
732,222 | 4.2924 ms |
antd.js |
6,683,633 | 16.9652 ms |
These figures come from one machine and are illustrative, not a regression baseline.
Results vary between runs, most noticeably for large fixtures such as antd.js.