Commit Graph

70 Commits

Author SHA1 Message Date
Dunqing cf4b7d7402 refactor(ast): migrate JavaScript string values to JSStr
JavaScript strings can contain lone UTF-16 surrogates. Keep their full
values through parsing, semantic analysis, transforms, linting, and
output instead of losing names at fallible UTF-8 views.

Build on the separate StaticName API preparation. Preserve canonical
WTF-8 through syntax consumers and convert explicitly at UTF-8 output,
identifier, pattern, and resolver boundaries. Keep enum constants in
owned UTF-16 storage without adding a general owned JSString.

Retain JSX character-reference output, generated CI files, and the
consumer regressions. The prerequisite isolates the property-name API
review and preserves allocation counts when formatting UTF-8 names.

Part of #26242.

Implemented with AI assistance.
2026-09-14 17:19:08 +08:00
camc314 8a9bdbd6d8 fix(estree): include decorators in FormalParameterRest spans (#26021)
Serialize `FormalParameterRest` with its outer span so its range includes parameter decorators. Regenerate the raw-transfer deserializers and remove `sourceMapValidationDecorators.ts` from the ESTree mismatch baseline.

Fixes #26011.
2026-08-23 18:32:30 +00:00
Boshen 1193c9dbaf chore(oxfmt): dogfood operator position at line start (#25937)
dogfood
2026-08-20 14:36:47 +00:00
camc314 0c68b7ff8b fix(estree): emit decorators on FormalParameterRest (#25582)
This PR adds support for emitting `decorators` when serializing an ast to ESTree (both with/without raw transfer).

Given the following code:
```
class C { method(@dec ...args) {} }
```

Before we emitted the following ESTree shape:
```json
"params": [
  {
    "type": "RestElement",
    "decorators": [],
    "argument": {
      "type": "Identifier",
      "decorators": [],
      "name": "args",
      "optional": false,
      "typeAnnotation": null,
    },
    "optional": false,
    "typeAnnotation": null,
    "value": null,
  }
],
```

Now we emit:
```json
"params": [
  {
    "type": "RestElement",
    "decorators": [
      {
        "type": "Decorator",
        "expression": { "type": "Identifier", "name": "dec" }
      }
    ],
    "argument": {
      "type": "Identifier",
      "decorators": [],
      "name": "args",
      "optional": false,
      "typeAnnotation": null,
    },
    "optional": false,
    "typeAnnotation": null,
    "value": null,
  },
],
```

fixes https://github.com/oxc-project/oxc/issues/18981
closes https://github.com/oxc-project/oxc/pull/24736
2026-08-13 10:43:14 +00:00
camc314 a33788e13e feat(ast)!: group class heritage into ClassHeritage (#25360)
## Summary

- Groups `Class::super_class` and `Class::super_type_arguments` into `Class::heritage`.
- Introduces `ClassHeritage` containing the superclass expression and its optional TypeScript type arguments.
- Makes superclass type arguments without a superclass expression unrepresentable.

## Breaking Change

The Rust AST changes from:

```rust
pub struct Class<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub r#type: ClassType,
    pub decorators: Vec<'a, Decorator<'a>>,
    pub id: Option<BindingIdentifier<'a>>,
    #[ts]
    pub type_parameters: Option<Box<'a, TSTypeParameterDeclaration<'a>>>,
    pub super_class: Option<Expression<'a>>,
    #[ts]
    pub super_type_arguments: Option<Box<'a, TSTypeParameterInstantiation<'a>>>,
    #[ts]
    pub implements: Vec<'a, TSClassImplements<'a>>,
    pub body: Box<'a, ClassBody<'a>>,
    #[ts]
    pub r#abstract: bool,
    #[ts]
    pub declare: bool,
    pub scope_id: Cell<Option<ScopeId>>,
}
```

to:

```rust
pub struct Class<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub r#type: ClassType,
    pub decorators: Vec<'a, Decorator<'a>>,
    pub id: Option<BindingIdentifier<'a>>,
    #[ts]
    pub type_parameters: Option<Box<'a, TSTypeParameterDeclaration<'a>>>,
    pub heritage: Option<ClassHeritage<'a>>,
    #[ts]
    pub implements: Vec<'a, TSClassImplements<'a>>,
    pub body: Box<'a, ClassBody<'a>>,
    #[ts]
    pub r#abstract: bool,
    #[ts]
    pub declare: bool,
    pub scope_id: Cell<Option<ScopeId>>,
}

pub struct ClassHeritage<'a> {
    pub expression: Expression<'a>,
    #[ts]
    pub type_arguments: Option<Box<'a, TSTypeParameterInstantiation<'a>>>,
}
```

## Why?

ECMAScript defines the complete `extends` clause as `ClassHeritage`:

```text
ClassHeritage : extends LeftHandSideExpression
```

TypeScript additionally permits type arguments on the heritage expression:

```ts
class Derived extends Base<T> {}
```

Previously, the AST stored the expression and type arguments independently:

```text
Class {
    super_class: Some(Base),
    super_type_arguments: Some(<T>),
    ...
}
```

This allowed an invalid state to be constructed:

```text
Class {
    super_class: None,
    super_type_arguments: Some(<T>),
    ...
}
```

Superclass type arguments cannot exist without an `extends` expression.

The parser already treated these fields as a pair. It parsed a class heritage entry and then split its expression and type arguments into two independent fields. Builders and transforms therefore had to preserve their relationship manually.

Grouping them into `ClassHeritage` makes the syntactic dependency explicit and prevents consumers from observing or constructing contradictory class heritage.

## Class heritage forms

A class without an `extends` clause has no heritage:

```js
class Base {}
```

```text
Class {
    heritage: None,
    ...
}
```

A class extending an expression has heritage without type arguments:

```js
class Derived extends Base {}
```

```text
Class {
    heritage: Some(ClassHeritage {
        expression: Base,
        type_arguments: None,
    }),
    ...
}
```

TypeScript type arguments belong to the same heritage structure:

```ts
class Derived extends Base<T> {}
```

```text
Class {
    heritage: Some(ClassHeritage {
        expression: Base,
        type_arguments: Some(<T>),
    }),
    ...
}
```

The heritage expression may be any valid class heritage expression:

```js
class Derived extends mixin(Base) {}
class NullDerived extends null {}
```

## How to migrate

Code that previously accessed the superclass expression directly:

```rust
if let Some(super_class) = &class.super_class {
    // ...
}
```

can use the compatibility accessor:

```rust
if let Some(super_class) = class.super_class() {
    // ...
}
```

Code that needs both parts should access the grouped heritage:

```rust
if let Some(heritage) = &class.heritage {
    let expression = &heritage.expression;
    let type_arguments = heritage.type_arguments.as_deref();
}
```

Mutating consumers should update the nested fields:

```rust
if let Some(heritage) = &mut class.heritage {
    heritage.type_arguments = None;
}
```

Common migrations:

| Previous representation | New representation |
| --- | --- |
| `class.super_class` | `class.heritage.as_ref().map(\|h\| &h.expression)` or `class.super_class()` |
| `class.super_type_arguments` | `class.heritage.as_ref()?.type_arguments` or `class.super_type_arguments()` |
| `class.super_class.is_some()` | `class.heritage.is_some()` |
| Mutate `class.super_class` | Mutate `class.heritage.expression` |
| Clear `class.super_type_arguments` | Clear `class.heritage.type_arguments` |
| Pass superclass and type arguments separately to a `Class` builder | Construct and pass `ClassHeritage` |
| Clone both fields independently | Clone `class.heritage` |

```ts
superClass: Expression | null;
superTypeArguments?: TSTypeParameterInstantiation | null;
```

No `heritage` property or `ClassHeritage` node is added to the ESTree output.

fixes https://github.com/oxc-project/backlog/issues/216
2026-08-07 14:18:59 +00:00
camc314 5c5cdcd522 feat(ast)!: narrow TSInterfaceHeritage::expression to TSTypeName (#24360)
## Summary

Narrow `TSInterfaceHeritage` from a general JavaScript `Expression` to `TSTypeName`.

```rust
pub struct TSInterfaceHeritage<'a> {
    pub type_name: TSTypeName<'a>,
    pub type_arguments: Option<Box<'a, TSTypeParameterInstantiation<'a>>>,
}
```

## Breaking Change

The Rust AST field changes from:

```rust
expression: Expression<'a>
```

to:

```rust
type_name: TSTypeName<'a>
```

## Why?

The current AST can represent shapes such as:

```ts
interface A extends foo() {}
interface A extends A + B {}
interface A extends new Foo() {}
interface A extends true {}
```

Which is an overly wide type since all of the above variants are invalid.

This PR tightens up the AST, making these invalid variants irrepresentable.

## How to migrate?

This should be a trivial migration:
1. instead of accessing `expression` on `TSInterfaceHeritage`, now access `type_name`
2. change any pattern matching on the `type_name` field to use `TypeName` (this is a smaller subset so should allow deleting code!)

closes https://github.com/oxc-project/backlog/issues/215
2026-08-07 06:00:42 +00:00
camc314 6be314f226 feat(ast)!: remove duplicated VariableDeclarator::kind (#25319)
## Summary

- Removes `VariableDeclarator::kind`.
- Makes `VariableDeclaration::kind` the single source of truth for every declarator in its declaration list.
- Makes contradictory parent and child declaration kinds unrepresentable.

## Breaking Change

The Rust AST changes from:

```rust
pub struct VariableDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub kind: VariableDeclarationKind,
    pub declarations: Vec<'a, VariableDeclarator<'a>>,
    #[ts]
    pub declare: bool,
}

pub struct VariableDeclarator<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    #[estree(skip)]
    pub kind: VariableDeclarationKind,
    #[estree(via = VariableDeclaratorId)]
    pub id: BindingPattern<'a>,
    #[ts]
    #[estree(skip)]
    pub type_annotation: Option<Box<'a, TSTypeAnnotation<'a>>>,
    pub init: Option<Expression<'a>>,
    #[ts]
    pub definite: bool,
}
```

to:

```rust
pub struct VariableDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub kind: VariableDeclarationKind,
    pub declarations: Vec<'a, VariableDeclarator<'a>>,
    #[ts]
    pub declare: bool,
}

pub struct VariableDeclarator<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    #[estree(via = VariableDeclaratorId)]
    pub id: BindingPattern<'a>,
    #[ts]
    #[estree(skip)]
    pub type_annotation: Option<Box<'a, TSTypeAnnotation<'a>>>,
    pub init: Option<Expression<'a>>,
    #[ts]
    pub definite: bool,
}
```

## Why?

A variable declaration has one declaration kind shared by every declarator in its declaration list:

```js
const first = 1, second = 2;
```

Previously, the AST stored that kind on both the parent declaration and every child declarator:

```text
VariableDeclaration {
    kind: Const,
    declarations: [
        VariableDeclarator { kind: Const, ... },
        VariableDeclarator { kind: Const, ... },
    ],
}
```

This allowed contradictory ASTs to be constructed:

```text
VariableDeclaration {
    kind: Const,
    declarations: [
        VariableDeclarator { kind: Var, ... },
    ],
}
```

Different consumers read different copies:

- Code generation used `VariableDeclaration::kind`.
- Semantic binding used `VariableDeclarator::kind`.
- Lint rules, transforms, and minifier passes used a mixture of both.

As a result, one AST could be emitted as `const` while its binding was treated as function-scoped `var`.

Transforms also had to keep the copies synchronized manually. For example, const-to-let normalization updated the parent and then looped through every declarator to update each child. Explicit resource management performed similar paired mutations for `using` declarations.

Keeping the kind only on `VariableDeclaration` removes this synchronization requirement and prevents consumers from observing conflicting declaration semantics.

## Declaration forms

The declaration kind now belongs exclusively to the node that owns the declaration list:

```js
let value;
```

```text
VariableDeclaration {
    kind: Let,
    declarations: [
        VariableDeclarator { id: value, ... },
    ],
}
```

Multiple declarators share the same parent kind:

```js
const first = 1, second = 2;
```

```text
VariableDeclaration {
    kind: Const,
    declarations: [
        VariableDeclarator { id: first, ... },
        VariableDeclarator { id: second, ... },
    ],
}
```

## How to migrate

Code that previously read the kind from a declarator:

```rust
if declarator.kind.is_const() {
    // ...
}
```

should now read it from the containing declaration:

```rust
if declaration.kind.is_const() {
    // ...
}
```

Consumers operating through an AST node store should retrieve the parent `VariableDeclaration`:

```rust
let AstKind::VariableDeclaration(declaration) =
    nodes.parent_kind(declarator.node_id())
else {
    unreachable!();
};

let kind = declaration.kind;
```

Common migrations:

| Previous representation | New representation |
| --- | --- |
| `declarator.kind` | Containing `declaration.kind` |
| Match `VariableDeclarator::kind` during semantic binding | Read the parent `VariableDeclaration` from ancestry |
| Pass `kind` to `VariableDeclarator::new` | Remove the `kind` argument |
| Update parent and every child kind | Update only `VariableDeclaration::kind` |
| Recover kind after detaching a declarator | Capture or pass the parent kind before detaching it |
| Inspect a symbol’s declarator kind in lint rules | Resolve its parent declaration through the node store |

When constructing declarators, use the generated kindless `VariableDeclarator` builders and place them inside a `VariableDeclaration` carrying the required kind.

fixes https://github.com/oxc-project/backlog/issues/237
2026-08-05 16:31:46 +00:00
camc314 44fd320324 feat(ast)!: split TS external modules & Namespace Declarations (#25284)
## Summary

- Splits `TSModuleDeclaration` into two dedicated declaration nodes:
  - `TSExternalModuleDeclaration` for string-literal modules and module augmentations.
  - `TSNamespaceDeclaration` for identifier-based `module` and `namespace` declarations.
- Removes `TSModuleDeclaration`, `TSModuleDeclarationName`, `TSModuleDeclarationBody`, and `TSModuleDeclarationKind`.
- Makes invalid combinations of names, bodies, and declaration kinds unrepresentable.
- Ensures only namespace declarations introduce a normal binding.
- Updates all AST consumers, generators, traversal, formatting, code generation, semantic analysis, lint rules, transforms, and serialization.
- Keeps normal and eager-raw ESTree output compatible: both declarations remain `TSModuleDeclaration` nodes.
- Allows experimental lazy raw deserialization to expose the two new node types directly.

## Breaking Change

The Rust AST changes from:

```rust
pub enum Declaration<'a> {
    VariableDeclaration(Box<'a, VariableDeclaration<'a>>) = 32,
    FunctionDeclaration(Box<'a, Function<'a>>) = 33,
    ClassDeclaration(Box<'a, Class<'a>>) = 34,

    TSTypeAliasDeclaration(Box<'a, TSTypeAliasDeclaration<'a>>) = 35,
    TSInterfaceDeclaration(Box<'a, TSInterfaceDeclaration<'a>>) = 36,
    TSEnumDeclaration(Box<'a, TSEnumDeclaration<'a>>) = 37,
    TSModuleDeclaration(Box<'a, TSModuleDeclaration<'a>>) = 38,
    TSGlobalDeclaration(Box<'a, TSGlobalDeclaration<'a>>) = 39,
    TSImportEqualsDeclaration(Box<'a, TSImportEqualsDeclaration<'a>>) = 40,
}

pub struct TSModuleDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub id: TSModuleDeclarationName<'a>,
    pub body: Option<TSModuleDeclarationBody<'a>>,
    pub kind: TSModuleDeclarationKind,
    pub declare: bool,
    pub scope_id: Cell<Option<ScopeId>>,
}
```

to:

```rust
pub enum Declaration<'a> {
    VariableDeclaration(Box<'a, VariableDeclaration<'a>>) = 32,
    FunctionDeclaration(Box<'a, Function<'a>>) = 33,
    ClassDeclaration(Box<'a, Class<'a>>) = 34,

    TSTypeAliasDeclaration(Box<'a, TSTypeAliasDeclaration<'a>>) = 35,
    TSInterfaceDeclaration(Box<'a, TSInterfaceDeclaration<'a>>) = 36,
    TSEnumDeclaration(Box<'a, TSEnumDeclaration<'a>>) = 37,
    TSExternalModuleDeclaration(Box<'a, TSExternalModuleDeclaration<'a>>) = 38,
    TSNamespaceDeclaration(Box<'a, TSNamespaceDeclaration<'a>>) = 39,
    TSGlobalDeclaration(Box<'a, TSGlobalDeclaration<'a>>) = 40,
    TSImportEqualsDeclaration(Box<'a, TSImportEqualsDeclaration<'a>>) = 41,
}

pub struct TSExternalModuleDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub id: StringLiteral<'a>,
    pub body: Option<Box<'a, TSModuleBlock<'a>>>,
    pub declare: bool,
    pub scope_id: Cell<Option<ScopeId>>,
}

pub struct TSNamespaceDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub id: BindingIdentifier<'a>,
    pub body: TSNamespaceDeclarationBody<'a>,
    pub kind: TSNamespaceDeclarationKind,
    pub declare: bool,
    pub scope_id: Cell<Option<ScopeId>>,
}

pub enum TSNamespaceDeclarationBody<'a> {
    TSNamespaceDeclaration(Box<'a, TSNamespaceDeclaration<'a>>) = 0,
    TSModuleBlock(Box<'a, TSModuleBlock<'a>>) = 1,
}

pub enum TSNamespaceDeclarationKind {
    Module = 0,
    Namespace = 1,
}
```

The corresponding inherited `Statement` variants, `AstKind` variants, builders, and visitor methods have also changed.

## Why?

`TSModuleDeclaration` previously represented two declaration families with different syntax and semantics:

```ts
declare module "foo" {}
declare module "*.css";

module Foo {}
namespace Foo {}
namespace Foo.Bar {}
```

The shared representation allowed invalid states such as:

- A namespace with a string-literal name.
- An external module with an identifier name.
- An identifier namespace without a body.
- A dotted external module.
- An external module containing another module declaration as its body.
- A declaration whose `kind` disagreed with its name.

Consumers also had to inspect `id` before determining whether the node introduced a binding or represented an external module.

The two declaration families have stronger independent invariants:

### External modules

- Always have a string-literal name.
- Never introduce a normal binding identifier.
- May omit their body.
- Cannot use dotted names.
- Use module-body import and export rules.

### Namespace declarations

- Always have a binding identifier.
- Always have a body.
- May contain nested namespace declarations for dotted names.
- Explicitly preserve the `module` or `namespace` spelling.
- Introduce a namespace/module symbol.
- Use namespace-body restrictions.

Dedicated nodes encode these differences directly.

## Declaration forms

External modules use `TSExternalModuleDeclaration`:

```ts
declare module "foo" {}
declare module "*.css";
module "foo" {}
```

Identifier-based declarations use `TSNamespaceDeclaration`:

```ts
module Foo {}
namespace Foo {}
```

Dotted namespaces are represented as nested namespace declarations:

```ts
namespace Foo.Bar {}
```

```rust
TSNamespaceDeclaration {
    id: Foo,
    body: TSNamespaceDeclarationBody::TSNamespaceDeclaration(
        TSNamespaceDeclaration {
            id: Bar,
            body: TSNamespaceDeclarationBody::TSModuleBlock(...),
            // ...
        }
    ),
    // ...
}
```

Only `TSExternalModuleDeclaration::body` is optional.

## ESTree compatibility

Normal serialization and eager raw deserialization continue to emit the existing ESTree shape:

```ts
interface TSModuleDeclaration {
  type: "TSModuleDeclaration";
  id: BindingIdentifier | StringLiteral | TSQualifiedName;
  body: TSModuleBlock | null;
  kind: "module" | "namespace";
  declare: boolean;
  global: false;
}
```

Dotted namespaces continue to serialize as a single `TSModuleDeclaration` with a `TSQualifiedName`.

Experimental lazy raw deserialization is not compatibility-preserving. It now exposes `TSExternalModuleDeclaration` and `TSNamespaceDeclaration` directly, with separate constructors, layouts, and visitor type IDs.

## How to migrate

Code that previously matched all forms through one variant:

```rust
if let Declaration::TSModuleDeclaration(declaration) = declaration {
    // ...
}
```

should now match the required declaration family:

```rust
match declaration {
    Declaration::TSExternalModuleDeclaration(declaration) => {
        // String-literal external module
    }
    Declaration::TSNamespaceDeclaration(declaration) => {
        // Identifier-based module or namespace
    }
    _ => {}
}
```

The same applies to `Statement` and `AstKind`:

```rust
Statement::TSExternalModuleDeclaration(declaration)
Statement::TSNamespaceDeclaration(declaration)

AstKind::TSExternalModuleDeclaration(declaration)
AstKind::TSNamespaceDeclaration(declaration)
```

Common field migrations:

| Previous representation | New representation |
| --- | --- |
| `TSModuleDeclarationName::StringLiteral(id)` | `TSExternalModuleDeclaration::id` |
| `TSModuleDeclarationName::Identifier(id)` | `TSNamespaceDeclaration::id` |
| `body: None` | Only valid on `TSExternalModuleDeclaration` |
| `TSModuleDeclarationBody::TSModuleBlock` | External `body` or `TSNamespaceDeclarationBody::TSModuleBlock` |
| `TSModuleDeclarationBody::TSModuleDeclaration` | `TSNamespaceDeclarationBody::TSNamespaceDeclaration` |
| `TSModuleDeclarationKind::Module` | External module, or `TSNamespaceDeclarationKind::Module` |
| `TSModuleDeclarationKind::Namespace` | `TSNamespaceDeclarationKind::Namespace` |
| Inspect `id` to determine binding behavior | Match `TSNamespaceDeclaration` |
| `declaration.id()` for string modules | Now returns `None` |
| `visit_ts_module_declaration` | `visit_ts_external_module_declaration` or `visit_ts_namespace_declaration` |

When constructing nodes, use the generated `TSExternalModuleDeclaration` or `TSNamespaceDeclaration` builders rather than recreating the previous shared representation.

fixes https://github.com/oxc-project/backlog/issues/218
2026-08-05 14:37:35 +00:00
camc314 067da8c4e7 feat(ast)!: store single parameter in TSIndexSignature::parameter (#25154) 2026-07-31 20:35:37 +00:00
camc314 1bdedd11ca feat(ast)!: introduce ExportDeclaration, ExportFromDeclaration (#25095)
## Summary

Split the overloaded Rust `ExportNamedDeclaration` into syntax-specific AST nodes:

- `ExportDeclaration` for exported declarations
- `ExportNamedDeclaration` for local exports
- `ExportFromDeclaration` for re-exports

## Breaking Change

The Rust AST changes from:

```rust
pub struct ExportNamedDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub declaration: Option<Declaration<'a>>,
    pub specifiers: Vec<'a, ExportSpecifier<'a>>,
    pub source: Option<StringLiteral<'a>>,
    pub export_kind: ImportOrExportKind,
    pub with_clause: Option<Box<'a, WithClause<'a>>>,
}

pub enum ModuleDeclaration<'a> {
    // ...
    ExportNamedDeclaration(Box<'a, ExportNamedDeclaration<'a>>),
    // ...
}
```

to:

```rust
pub struct ExportDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub declaration: Declaration<'a>,
}

pub struct ExportNamedDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub specifiers: Vec<'a, ExportSpecifier<'a>>,
    pub export_kind: ImportOrExportKind,
}

pub struct ExportFromDeclaration<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub specifiers: Vec<'a, ExportSpecifier<'a>>,
    pub source: StringLiteral<'a>,
    pub export_kind: ImportOrExportKind,
    pub with_clause: Option<Box<'a, WithClause<'a>>>,
}

pub enum ModuleDeclaration<'a> {
    // ...
    ExportDeclaration(Box<'a, ExportDeclaration<'a>>),
    ExportNamedDeclaration(Box<'a, ExportNamedDeclaration<'a>>),
    ExportFromDeclaration(Box<'a, ExportFromDeclaration<'a>>),
    // ...
}
```

## Why?

The previous node represented three distinct syntax forms through optional and mutually exclusive fields:

```text
export const value = 1;       -> ExportDeclaration
export { value };             -> ExportNamedDeclaration
export { value } from "mod";  -> ExportFromDeclaration
```

This allowed invalid combinations such as declarations with specifiers or sources, and import attributes without a source. The new nodes make these states unrepresentable and remove repeated `declaration`/`source` branching from consumers.

## How to migrate

Match the syntax-specific `ModuleDeclaration`, `Statement`, or `AstKind` variant instead of inspecting optional fields. Generated builders, visitors, traversal ancestors, and formatter wrappers now expose corresponding APIs for all three nodes.

Common field migrations:

| Previous representation | New representation |
| --- | --- |
| `declaration: Some(declaration)` | `ExportDeclaration { declaration }` |
| `declaration.as_ref()` | `&export.declaration` |
| `source.is_none()` for local exports | Match `ExportNamedDeclaration` |
| `source: Some(source)` | `ExportFromDeclaration { source }` |
| `source.as_ref().unwrap()` | `&export.source` |
| `with_clause` | Only available on `ExportFromDeclaration` |
| Stored `export_kind` on exported declarations | `ExportDeclaration::export_kind()` |
| `ExportNamedDeclaration::boxed_plain_declaration(...)` | `ExportDeclaration::boxed(...)` |
| `ExportNamedDeclaration::boxed(..., None, specifiers, None, ...)` | `ExportNamedDeclaration::boxed(..., specifiers, ...)` |
| `ExportNamedDeclaration::boxed(..., None, specifiers, Some(source), ...)` | `ExportFromDeclaration::boxed(..., specifiers, source, ...)` |

closes https://github.com/oxc-project/backlog/issues/219
2026-07-31 10:25:47 +00:00
camc314 c917f204a7 feat(ast)!: introduce ArrowFunctionBody enum (#24987)
## Summary

- Introduces `ArrowFunctionBody` to represent concise expression and block bodies explicitly.
- Removes the `ArrowFunctionExpression::expression` boolean.
- Makes invalid arrow body combinations unrepresentable.

## Breaking Change

The Rust AST changes from:

```rust
pub struct ArrowFunctionExpression<'a> {
    pub expression: bool,
    pub body: Box<'a, FunctionBody<'a>>,
}
```

to:

```rust
pub enum ArrowFunctionBody<'a> {
    FunctionBody(Box<'a, FunctionBody<'a>>) = 64,
    INHERIT(Expression<'a>),
}

pub struct ArrowFunctionExpression<'a> {
    pub body: ArrowFunctionBody<'a>,
}
```

Concise arrow bodies are now stored directly as expressions instead of synthetic `FunctionBody` and `ExpressionStatement` nodes. Block bodies continue to use `FunctionBody`.

## Why?

The previous representation required consumers to maintain this implicit invariant:

```text
expression == true
→ no directives
→ exactly one statement
→ that statement is an ExpressionStatement
```

Invalid combinations could cause panics, skipped work, or fallback behavior. The enum encodes the grammar directly and makes those invalid states unrepresentable.

Concise expressions also now have `ArrowFunctionExpression` as their direct traversal ancestor, without synthetic `FunctionBody` or `ExpressionStatement` nodes.

## How to migrate

Use the arrow body accessors or match on `ArrowFunctionBody` rather than inspecting `expression` and synthetic statements.

Common field migrations:

| Previous representation | New representation |
| --- | --- |
| `arrow.expression` | `arrow.is_expression()` |
| `arrow.get_expression()` | `arrow.get_expression()` |
| `&arrow.body.statements` | `arrow.get_function_body().map(\|body\| &body.statements)` |
| `&mut arrow.body.statements` | `arrow.get_function_body_mut().map(\|body\| &mut body.statements)` |
| Synthetic expression statement in `arrow.body` | Direct `ArrowFunctionBody` expression variant |
| `arrow.body = function_body` | `arrow.body = ArrowFunctionBody::FunctionBody(function_body)` |
| Expression-to-block conversion | Create a `FunctionBody` containing a return statement |
| Block-to-expression conversion | Assign the expression directly to `arrow.body` |
| Builder `(expression, body)` arguments | Pass a single `ArrowFunctionBody` |

closes https://github.com/oxc-project/backlog/issues/211
2026-07-29 13:32:32 +00:00
overlookmotel e80574fdf0 fix(estree): handle empty spans serializing ImportMeta and NewTarget (#24775)
Follow-on after #24557. Handle when `ImportMeta` / `NewTarget` has an empty span (`SPAN`). This code path is not exercised at present, but it seems a good idea to be prepared for when we may serialize AST after transformation.
2026-07-22 08:05:57 +00:00
Armano 7b045cd417 feat(minfier): drop last break from last switch case (#24673)
After many trial and errors, this seem to be the best way to remove
break stmts in switch cases,

This change targets only last break in last case of switch smts if its
unlabelled

this change differs from #23914 as its no longer targets all branches as
that should not be no longer necessary if all cases are correctly
rewritten

https://github.com/oxc-project/oxc/issues/24672

-----

cases like `switch (a) { case 3: if (b) break; }` are not changed

-----

ref: #17544
2026-07-20 13:21:03 +08:00
camc314 54cc121250 feat(ast)!: split MetaProperty into ImportMeta and NewTarget (#24557)
## Summary
- replace the generic Rust MetaProperty AST node with dedicated ImportMeta and NewTarget nodes
- preserve the ESTree MetaProperty shape and generated public TypeScript types
- update parser, formatter, codegen, lint, transform, minifier, traversal, and React compiler consumers

## Summary

Remove `MetaProperty` AST node, in favour of two separate nodes (`ImportMeta` and `NewTarget`)

## Breaking Change

The Rust AST field changes from:

```rust
pub struct MetaProperty<'a> {
    pub node_id: Cell<NodeId>,
    pub span: Span,
    pub meta: IdentifierName<'a>,
    pub property: IdentifierName<'a>,
}
```

to:

```rust
/// /// `import.meta` in `console.log(import.meta);`.
pub struct ImportMeta {
    pub node_id: Cell<NodeId>,
    pub span: Span,
}

/// `new.target` in `function F() { return new.target; }`.
pub struct NewTarget {
    pub node_id: Cell<NodeId>,
    pub span: Span,
}
```

## Why?

The current AST could represent shapes such as:

```
foo.bar
```

By constructing the node manually. However this doesn't make sense, and is invalid.

## How to migrate?

1. Replace any `Expression::MetaProperty` matches/checks with the dedicated `Expression::ImportMeta` and `Expression::NewTarget` varients
2. Replace any `new_meta_property` construction with the new `new_import_meta` or `new_new_target` APIs.
3. Inspecting the `meta`/`property` identifier names is no longer required.
2026-07-15 14:48:32 +00:00
renovate[bot] 0e00b91b5a chore(deps): update dependency rust to v1.97.0 (#24328)
This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [rust](https://redirect.github.com/rust-lang/rust) | minor | `1.96.1`
→ `1.97.0` |

---

### Release Notes

<details>
<summary>rust-lang/rust (rust)</summary>

###
[`v1.97.0`](https://redirect.github.com/rust-lang/rust/blob/HEAD/RELEASES.md#Version-1970-2026-07-09)

[Compare
Source](https://redirect.github.com/rust-lang/rust/compare/1.96.1...1.97.0)

\==========================

<a id="1.97.0-Language"></a>

## Language

- [Consider `Result<T, Uninhabited>` and `ControlFlow<Uninhabited, T>`
to be equivalent to `T` for must use
lint](https://redirect.github.com/rust-lang/rust/pull/148214)
- [Add allow-by-default `dead_code_pub_in_binary` lint for unused pub
items in binary
crates](https://redirect.github.com/rust-lang/rust/pull/149509)
- [Stabilize the `div32`, `lam-bh`, `lamcas`, `ld-seq-sa` and `scq`
target features](https://redirect.github.com/rust-lang/rust/pull/154510)
- [Stabilize
`cfg(target_has_atomic_primitive_alignment)`](https://redirect.github.com/rust-lang/rust/pull/155006)
- [Allow trailing `self` in imports in more
cases](https://redirect.github.com/rust-lang/rust/pull/155137)

<a id="1.97.0-Platform-Support"></a>

## Platform Support

- [nvptx64-nvidia-cuda: drop support for old architectures and old
ISAs](https://redirect.github.com/rust-lang/rust/pull/152443)

Refer to Rust's [platform support page][platform-support-doc]
for more information on Rust's tiered platform support.

[platform-support-doc]:
https://doc.rust-lang.org/rustc/platform-support.html

<a id="1.97.0-Stabilized-APIs"></a>

## Stabilized APIs

- [`Default for
RepeatN`](https://doc.rust-lang.org/stable/std/iter/struct.RepeatN.html#impl-Default-for-RepeatN%3CA%3E)
- [`Copy for
ffi::FromBytesUntilNulError`](https://doc.rust-lang.org/stable/std/ffi/struct.FromBytesUntilNulError.html#impl-Copy-for-FromBytesUntilNulError)
- [`Send for std::fs::File` on
UEFI](https://redirect.github.com/rust-lang/rust/pull/154003)
-
[`<{integer}>::isolate_highest_one`](https://doc.rust-lang.org/stable/std/primitive.u32.html#method.isolate_highest_one)
-
[`<{integer}>::isolate_lowest_one`](https://doc.rust-lang.org/stable/std/primitive.u32.html#method.isolate_lowest_one)
-
[`<{integer}>::highest_one`](https://doc.rust-lang.org/stable/std/primitive.u32.html#method.highest_one)
-
[`<{integer}>::lowest_one`](https://doc.rust-lang.org/stable/std/primitive.u32.html#method.lowest_one)
-
[`<{integer}>::bit_width`](https://doc.rust-lang.org/stable/std/primitive.u32.html#method.bit_width)
-
[`NonZero<{integer}>::isolate_highest_one`](https://doc.rust-lang.org/stable/std/num/struct.NonZero.html#method.isolate_highest_one)
-
[`NonZero<{integer}>::isolate_lowest_one`](https://doc.rust-lang.org/stable/std/num/struct.NonZero.html#method.isolate_lowest_one)
-
[`NonZero<{integer}>::highest_one`](https://doc.rust-lang.org/stable/std/num/struct.NonZero.html#method.highest_one)
-
[`NonZero<{integer}>::lowest_one`](https://doc.rust-lang.org/stable/std/num/struct.NonZero.html#method.lowest_one)
-
[`NonZero<{integer}>::bit_width`](https://doc.rust-lang.org/stable/std/num/struct.NonZero.html#method.bit_width)

These previously stable APIs are now stable in const contexts:

-
[`char::is_control`](https://doc.rust-lang.org/stable/std/primitive.char.html#method.is_control)

<a id="1.97.0-Cargo"></a>

## Cargo

- [Stabilize `build.warnings`
config.](https://redirect.github.com/rust-lang/cargo/pull/16796) This
controls how lint warnings from local packages are treated. Useful for
enforcing a warning-free build in CI, replacing `-Dwarnings`.
[docs](https://doc.rust-lang.org/nightly/cargo/reference/config.html#buildwarnings)
- [Stabilize `resolver.lockfile-path`
config.](https://redirect.github.com/rust-lang/cargo/pull/16694) This
allows specifying the path to the lockfile to use when resolving
dependencies. Useful when working with read-only source directories.
[docs](https://doc.rust-lang.org/nightly/cargo/reference/config.html#resolverlockfile-path)
- [cargo-clean: Error when `--target-dir` doesn't look like a Cargo
target
directory.](https://redirect.github.com/rust-lang/cargo/pull/16712) This
prevents accidental deletion of non-target directories.
- [Add `-m` shorthand for
`--manifest-path`](https://redirect.github.com/rust-lang/cargo/pull/16858)
- [Remove `curl` dependency from `crates-io`
crate](https://redirect.github.com/rust-lang/cargo/pull/16936)

<a id="1.97.0-Rustdoc"></a>

## Rustdoc

- [Stabilize `--emit`
flag](https://redirect.github.com/rust-lang/rust/pull/146220)
- [Stabilize
`--remap-path-prefix`](https://redirect.github.com/rust-lang/rust/pull/155307)

<a id="1.97.0-Compatibility-Notes"></a>

## Compatibility Notes

- [Emit a future-compatibility warning when relying on `f32:
From<{float}>` to constrain
`{float}`](https://redirect.github.com/rust-lang/rust/pull/139087)
- [Rust will use the v0 symbol mangling scheme by
default.](https://redirect.github.com/rust-lang/rust/pull/151994) This
may cause some tools (such as debuggers or profilers, especially with
old versions) to fail to demangle symbols emitted by Rust. It may also
cause the formatting of text in backtraces to change.
- [Prevent deref coercions in `pin!`, in order to prevent
unsoundness.](https://redirect.github.com/rust-lang/rust/pull/153457)
The most likely case where this might impact users is: writing `pin!(x)`
where `x` has type `&mut T` will now always correctly produce a value of
type `Pin<&mut &mut T>`, instead of sometimes allowing a coercion that
produces a value of type `Pin<&mut T>`. This coercion was previously
incorrectly allowed since Rust 1.88.0.
- [Deprecate `std::char` constants and
functions](https://redirect.github.com/rust-lang/rust/pull/153873)
- [Warn on linker output by
default](https://redirect.github.com/rust-lang/rust/pull/153968)
- [Remove hidden `f64` methods which have been deprecated since
1.0](https://redirect.github.com/rust-lang/rust/pull/153975)
- [report the `varargs_without_pattern` lint in
deps](https://redirect.github.com/rust-lang/rust/pull/154599)
- [Forbid passing generic arguments to module path segments even if the
module reexports a generic enum
variant](https://redirect.github.com/rust-lang/rust/pull/154971)
- [Error on invalid macho `link_section`
specifier](https://redirect.github.com/rust-lang/rust/pull/155065)
- The encoding of certain `enum`s [have
changed](https://redirect.github.com/rust-lang/rust/pull/155473). This
is not a breaking change, as it only applies to `enum`s without layout
guarantees, but is noted here as we've seen people impacted from having
made assumptions about the layout algorithm.
- [Error on `#[export_name = "..."]` where the name is
empty](https://redirect.github.com/rust-lang/rust/pull/155515)
- [Syntactically reject tuple index shorthands in struct
patterns](https://redirect.github.com/rust-lang/rust/pull/155698)
- [validate `#[link_name = "..."]` & `#[link(name = "...")]`
parameters](https://redirect.github.com/rust-lang/rust/pull/155817)
- On Windows, after calling `shutdown` on a socket to shut down the
write side, attempting to write to the socket will now produce a
`BrokenPipe` error rather than `Other`. [Map `WSAESHUTDOWN` to
`io::ErrorKind::BrokenPipe`](https://redirect.github.com/rust-lang/rust/pull/156063)

</details>

---

### Configuration

📅 **Schedule**: (in timezone Asia/Shanghai)

- Branch creation
  - At any time (no schedule defined)
- Automerge
  - At any time (no schedule defined)

🚦 **Automerge**: Enabled.

♻ **Rebasing**: Whenever PR is behind base branch, or you tick the
rebase/retry checkbox.

🔕 **Ignore**: Close this PR and you won't be reminded about this update
again.

---

- [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check
this box

---

This PR was generated by [Mend Renovate](https://mend.io/renovate/).
View the [repository job
log](https://developer.mend.io/github/oxc-project/oxc).

<!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4yNDIuMiIsInVwZGF0ZWRJblZlciI6IjQzLjI0Mi4yIiwidGFyZ2V0QnJhbmNoIjoibWFpbiIsImxhYmVscyI6W119-->

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: Cameron <cameron.clark@hey.com>
2026-07-09 18:34:23 +01:00
renovate[bot] 81511b4568 chore(deps): update dependency oxfmt to ^0.58.0 (#24240)
This PR contains the following updates:

| Package | Change |
[Age](https://docs.renovatebot.com/merge-confidence/) |
[Adoption](https://docs.renovatebot.com/merge-confidence/) |
[Passing](https://docs.renovatebot.com/merge-confidence/) |
[Confidence](https://docs.renovatebot.com/merge-confidence/) |
|---|---|---|---|---|---|
| [oxfmt](https://oxc.rs/docs/guide/usage/formatter)
([source](https://redirect.github.com/oxc-project/oxc/tree/HEAD/npm/oxfmt))
| [`^0.57.0` →
`^0.58.0`](https://renovatebot.com/diffs/npm/oxfmt/0.57.0/0.58.0) |
![age](https://developer.mend.io/api/mc/badges/age/npm/oxfmt/0.58.0?slim=true)
|
![adoption](https://developer.mend.io/api/mc/badges/adoption/npm/oxfmt/0.58.0?slim=true)
|
![passing](https://developer.mend.io/api/mc/badges/compatibility/npm/oxfmt/0.57.0/0.58.0?slim=true)
|
![confidence](https://developer.mend.io/api/mc/badges/confidence/npm/oxfmt/0.57.0/0.58.0?slim=true)
|

---

### Release Notes

<details>
<summary>oxc-project/oxc (oxfmt)</summary>

###
[`v0.58.0`](https://redirect.github.com/oxc-project/oxc/compare/oxfmt_v0.57.0...oxfmt_v0.58.0)

[Compare
Source](https://redirect.github.com/oxc-project/oxc/compare/oxfmt_v0.57.0...oxfmt_v0.58.0)

</details>

---

### Configuration

📅 **Schedule**: (in timezone Asia/Shanghai)

- Branch creation
  - At any time (no schedule defined)
- Automerge
  - At any time (no schedule defined)

🚦 **Automerge**: Enabled.

♻ **Rebasing**: Whenever PR is behind base branch, or you tick the
rebase/retry checkbox.

🔕 **Ignore**: Close this PR and you won't be reminded about this update
again.

---

- [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check
this box

---

This PR was generated by [Mend Renovate](https://mend.io/renovate/).
View the [repository job
log](https://developer.mend.io/github/oxc-project/oxc).

<!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4yNDIuMiIsInVwZGF0ZWRJblZlciI6IjQzLjI0Mi4yIiwidGFyZ2V0QnJhbmNoIjoibWFpbiIsImxhYmVscyI6W119-->

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
Co-authored-by: Cameron <cameron.clark@hey.com>
2026-07-07 10:23:12 +01:00
Dunqing 53509a85df feat(minifier): treeshake pure typed arrays and Set/Map array literals (#23469)
## Summary

Mark side-effect-free `new` constructions pure so they are dropped when unused (and annotated for downstream tree-shaking):

- **TypedArrays** with a non-negative numeric-literal length (`new Int8Array(8)`, `new Uint8Array(1024)`). The buffer allocation runs no user code; a valid length is pure and a too-large length only throws a maximum-length `RangeError`, which the minifier is allowed to drop (see `docs/ASSUMPTIONS.md`). `new Int8Array(-1)` (negative-length `RangeError`), `new Int8Array(0n)` (BigInt `TypeError`), object, and non-literal arguments are kept.
- **`new Set([...])` / `new Map([[..], ..])`** with array-literal arguments; element side effects are preserved when the construction is dropped. `WeakSet`/`WeakMap` are excluded (their keys must be objects, so `new WeakSet([1])` throws), and `Map` entries must be array literals (`new Map([1])` throws).

This matches esbuild's per-constructor handling — the only other minifier that keeps every throwing form; Rollup, Terser (`unsafe`), and SWC drop some of these incorrectly. Cross-verified against the ES spec, esbuild, Terser, Rollup, and SWC.

Minified-size impact: react.development.js −0.04 kB; size-neutral elsewhere (real code uses dynamic buffer lengths, not numeric literals).
2026-06-16 07:46:28 +00:00
overlookmotel c2c8f806f9 refactor(napi/parser): raw transfer store source text at end of buffer (#22392)
Oxlint's raw transfer implementation stores the source text at end of the buffer. Prior to this PR `napi/parser` stored source text at the start of the buffer.

Bring `napi/parser` into line with Oxlint, by also storing source text at end of the buffer.

Making all versions of raw transfer that we use align with each other removes complication (all versions now have the same implementation of `deserializeStr`) and also removes some fragile code where it'd be easy to accidentally trigger UB (see the updated comments in `arena/fixed_size/windows.rs` and `arena/alloc_impl.rs`).

Unfortunately, storing source at end of buffer is slightly less efficient than storing it at the start. We'll switch all implementations to storing source at start of buffer when we move to storing all strings (source text and all other strings) together in a single contiguous block. But in meantime, this slight inefficiency is outweighed by the gain in safety.
2026-05-14 22:26:41 +00:00
overlookmotel 99eef72e2d refactor(allocator, linter/plugins, napi/parser): store raw transfer metadata within Arena chunk (#21869)
Large change to how raw transfer stores metadata about the allocations it uses.

Previously the two metadata structures `RawTransferMetadata` and `FixedSizeAllocator` after the chunk's `ChunkFooter`.

This had a few problems:

1. It was confusing and unwieldy - a recipe for bugs.
2. `ChunkFooter`'s memory was exposed on JS side - a bit unsafe as altering bytes in this region could easily trigger UB.
3. It entwined `Arena` (which is just the thing that allocates) with the details of exactly _what_ raw transfer allocates.
4. Imposed annoying alignment requirements, because `ChunkFooter` must be aligned on 16, and so anything after it must ensure it doesn't break that invariant.

Old layout:

```
                                                        WHOLE BLOCK - aligned on 4 GiB
<-----------------------------------------------------> Allocated block (`BLOCK_SIZE` bytes)

                                                        ALLOCATOR
<----------------------------------------->             `Allocator` chunk (`CHUNK_SIZE` bytes)
                                     <---->             `ChunkFooter` (aligned on 16)
<----------------------------------->                   `Allocator` chunk data storage (for AST)
                                                        (`ACTIVE_SIZE` bytes)

                                                        METADATA
                                           <---->       `RawTransferMetadata`
                                                 <----> `FixedSizeAllocatorMetadata`

                                                        BUFFER SENT TO JS
<----------------------------------------------->       Buffer sent to JS (`BUFFER_SIZE` bytes)
```

This PR moves `RawTransferMetadata` and `FixedSizeAllocatorMetadata` into the chunk itself. New layout:

```
                                                        WHOLE BLOCK - size 2 GiB - 16, aligned on 4 GiB
<-----------------------------------------------------> Allocated block (`BLOCK_SIZE` bytes)

                                                        ARENA
<-----------------------------------------------------> Chunk (fills whole block)
<-------------------------------------->                Allocatable region for AST (`ACTIVE_SIZE` bytes)
                                        <--->           `RawTransferMetadata`
                                             <--->      `FixedSizeAllocatorMetadata`
                                                  <---> `ChunkFooter` (aligned on 16, last in block)

                                                        BUFFER SENT TO JS
<------------------------------------------->           Buffer sent to JS (`BUFFER_SIZE` bytes)
```

`FixedSizeAllocatorMetadata` and `ChunkFooter` are no longer in the region which is shared with JS side. As far as `Arena` is concerned, they're now just some data (like any other data) which is allocated in the arena.

Also:

- Introduce more consistency to the naming of constants which specify the size and position of these various data structures in the arena.
- Add more const assertions to ensure everything is laid out and aligned as it should be.
2026-04-30 01:27:30 +00:00
overlookmotel 502e804ffa perf(ast)!: reduce size of TSTypePredicateName (#21711)
#21653 revealed that `TSTypePredicateName` has an unboxed variant `This`.

Before the introduction of `NodeId` fields, this was fine as `TSThisType` contained only a `Span` (8 bytes). But with `NodeId` field added, it's 16 bytes, which makes `TSTypePredicateName` larger too - 24 bytes.

Box `TSThisType` in this variant, bringing `TSTypePredicateName` back down to 16 bytes.

This also has the side-effect of making `TSTypePredicateName`'s `node_id`, `span`, and `span_mut` methods branchless.
2026-04-24 13:03:50 +00:00
overlookmotel 5651539aba perf(ast)!: reduce size of JSXExpression (#21710)
#21653 revealed that `JSXExpression` has an unboxed variant `EmptyExpression`.

Before the introduction of `NodeId` fields, this was fine as `JSXEmptyExpression` contained only a `Span` (8 bytes). But with `NodeId` field added, it's 16 bytes, which makes `JSXExpression` larger too - 24 bytes.

Box `JSXEmptyExpression` in this variant, bringing `JSXExpression` back down to 16 bytes.

This also has the side-effect of making `JSXExpression`'s `node_id`, `span`, and `span_mut` methods branchless.
2026-04-24 13:03:50 +00:00
overlookmotel c44e28068e perf(ast)!: reduce size of ArrayExpressionElement (#21709)
#21653 revealed that `ArrayExpressionElement` has an unboxed variant `Elision`.

Before the introduction of `NodeId` fields, this was fine as `Elison` contained only a `Span` (8 bytes). But with `NodeId` field added, it's 16 bytes, which makes `ArrayExpressionElement` larger too - 24 bytes.

Box `Elison` in this variant, bringing `ArrayExpressionElement` back down to 16 bytes.

This also has the side-effect of making `ArrayExpressionElement`'s `node_id`, `span`, and `span_mut` methods branchless.
2026-04-24 13:03:50 +00:00
bab 00fc13620d fix(codegen): preserve coverage comments before object properties (#21312)
Closes #21302

---------

Co-authored-by: Dunqing <dengqing0821@gmail.com>
2026-04-20 12:51:01 +08:00
overlookmotel addcd02d97 perf(napi/parser, linter/plugins): raw transfer deserializer for Vecs use shift instead of multiply where possible (#21142)
Tiny optimization to raw transfer. Use shift instead of multiply where possible when calculating the end position of a `Vec`'s contents.
2026-04-07 23:32:03 +00:00
overlookmotel 3068ded6a1 perf(napi/parser, linter/plugins): shift before add when calculating positions in raw transfer deserializer (#21141)
Small perf optimization to raw transfer. When calculating a `u32` offset, shift first then add (`(pos >> 2) + 4`) instead of adding first (`(pos + 16) >> 2`). If a function contains multiple `pos >> 2` calculations, V8 may be able to combine them into a single calculation. Previously it couldn't prove that `pos` is a multiple of 4, so couldn't make this optimization.
2026-04-07 23:32:03 +00:00
overlookmotel eb400b8bb2 perf(napi/parser, linter/plugins): remove uint32 buffer view (#21140)
Continuation of #21132. All uses of the `Uint32Array` view of the buffer in raw transfer code have been removed in preceding PRs. So we can now remove it entirely.
2026-04-07 23:14:37 +00:00
overlookmotel 26750857d5 perf(napi/parser): lazy deserialization use only Int32Array (#21139)
Continuation of #21132. Use `Int32Array` view of the buffer for all code in lazy deserialization. Lazy deserialization is not currently in use, but keeping it in line with the main raw transfer deserializer allows removing the `Unint32Array` view of the buffer entirely in #21140 without breaking the lazy deserializer tests.
2026-04-07 23:14:36 +00:00
overlookmotel 7a866138d1 perf(linter/plugins): use Int32Arrays for tokens and comments buffers (#21136)
Continuation of #21132. Use `Int32Array` view of the buffer for all operations related to tokens and comments.
2026-04-07 23:08:00 +00:00
overlookmotel 8c51121e4c perf(napi/parser, linter/plugins): raw transfer deserialize Span fields as i32s (#21135)
Continuation of #21132. Get span `start` and `end` from the `Int32Array` view of the buffer, instead of the `Uint32Array`view. This allows V8 to statically see that they are SMIs, and skip type checks.
2026-04-07 17:25:42 +00:00
overlookmotel bc1bcdd95e perf(napi/parser, linter/plugins): inline trivial raw transfer field deserializers into node object definitions (#21134)
Utilize the `#[estree(raw_deser_inline)]` attribute introduced in #21134 to mark all custom deserializers which do not depend on `parent` as inline-able.

This reduces deserializer code size, as well as being more performant in some cases, by avoiding write boundary checks.
2026-04-07 17:25:42 +00:00
overlookmotel c0278abf55 perf(napi/parser, linter/plugins): use Int32Array in raw transfer deserializer (#21132)
Use `int32` (`Int32Array`) instead of `uint32`(`Uint32Array`) for getting data from buffer in raw transfer deserializer. The buffer is 2 GiB in size, so all offsets within the buffer can be represented as a `u31` which can be stored in an `i32` without any loss, and with no risk of being interpreted as negative numbers.

Same as in #21129, the advantage of fetching offsets from an `Int32Array` is that V8 statically knows the value can be stored in an SMI, its native integer type, without any range checks.
2026-04-07 17:25:41 +00:00
overlookmotel c70a8e9853 refactor(napi/parser, linter/plugins): add int32 buffer view (#21131)
Add an `Int32Array` view of the buffer for raw transfer. It will be used in PRs later in this stack.
2026-04-07 17:25:41 +00:00
overlookmotel a4ac3ce514 refactor(linter/plugins): import getNodeLoc instead of injecting it (#21026)
Refacfor. In a strange pattern, we were injecting `getNodeLoc` into `deserialize.js` module on every call to `deserializeProgramOnly`. It's a static function, so just import it.
2026-04-03 22:05:22 +00:00
overlookmotel fb52383874 perf(napi/parser, linter/plugins): clear buffers and source texts earlier (#21025)
There's no need to defer clearing vars containing buffers and source text in raw transfer deserializer module until end of linting. It was a left-over from when this module contained code which lazily deserialized comments, but that's now done elsewhere.

Clear these vars as soon as deserialization is complete. Buffer and `sourceText` are still held in vars in other modules until linting completes, but `sourceTextLatin` (introduced in #21021) can be freed immediately.
2026-04-03 22:05:22 +00:00
overlookmotel 3b7dec4dbb perf(napi/parser, linter/plugins): use utf8Slice for decoding UTF-8 strings (#21022)
Optimize raw transfer string decoding.

Replace `TextDecoder("utf8")` with `Buffer.prototype.utf8Slice` to decode Rust UTF-8 strings to JS strings.

Benchmarks show this is an average 5% speed-up in `deserializeStr`, though the reason why is unclear. It skips various checks, and avoids creating a temporary `Uint8Array` (which is required with `TextDecoder.decode`) which trims 60 bytes of the size of the assembly created for `deserializeStr`. But that's not enough to explain the speed-up 🤷. Anyway, take the win.
2026-04-03 22:05:22 +00:00
overlookmotel 012c924a04 perf(napi/parser, linter/plugins): speed up decoding strings in raw transfer (#21021)
Improve perf of deserializing strings in raw transfer. This PR combines several optimizations, which have been tested and benchmarked in https://github.com/overlookmotel/oxc-raw-str-bench. This PR implements the version "latin-slice-onebyte64" from that repo, which is the current winner.

String deserialization is the main bottleneck in raw transfer, so speeding it up will likely make a large impact on deserialization overall.

This work follows on from #20834 which produced a major speed-up in many files by making files which contain some non-ASCII characters take the fast path of slicing `sourceText` more often.

This PR tackles the remainder - speeding up the fallback path where the fast path can't be taken.

## Optimizations

The optimizations in this PR are:

### Latin1

When source is not 100% ASCII, decode source text from buffer as Latin1.

A Latin1-decoded string represents each UTF-8 byte as a single Latin1 character, so it can be indexed into using UTF-8 offsets.

So when we can't slice the string from `sourceText` because the UTF-8 and UTF-16 offsets differ (after any non-ASCII character), loop through the string's bytes and check if they're all ASCII. If they are, the string can be sliced from `sourceTextLatin` instead, with the original UTF-8 offsets.

This is way faster than calling `textDecoder.decode`, as it avoids a call into C++. [Benchmarks show](https://github.com/overlookmotel/oxc-raw-str-bench/blob/4f96275efa9a35d5d27615abb27f21a137149cc0/README.md#apply30-vs-latin-vs-latin-source64) speed up of 55% on average, and up to 70% on some files.

### Latin1 decoding method

It turns out that `new TextDecoder("latin1").decode(arr)` doesn't actually decode to Latin1!

Per the WHATWG Encoding Standard, "latin1" is mapped to "windows-1252".

The result is that with `TextDecoder("latin1")`:

1. `decode` is quite complicated, requiring a 2-pass scan of the bytes to determine if they're all ASCII, followed by a 2nd pass to do the actual `windows-1252` decoding. If the string *does* contain any non-ASCII characters (which it always does in our usecase), NodeJS implements the decoding in JS, not native code. Slow.
2. `decode` produces a 2-byte-per-char string (`TWO_BYTE` in V8), which takes more memory, and is slower for all operations on it e.g. string comparison, hashing for use as an object key etc.

Instead, use `Buffer.prototype.latin1Slice` which:

1. Does a pure Latin1 decode, which is just a single `memcpy` call.
2. Produces a 1-byte-per-char string (`ONE_BYTE` in V8).

`latin1Slice` involves a call into C++, but we only do it once per file, so this cost is tiny in context of deserializing the whole AST.

### Latin1 string slicing

In the fast path, slice from the Latin1-decoded string, instead of `sourceText`. In the fast path, we know that all bytes of source comprising the string are ASCII, so no further checks are required.

This makes no difference on benchmarks for `deserializeStr` itself, but it may have beneficial effects downstream for code (e.g. lint rules) which access strings in the AST, e.g. `Identifier` names.

Because Latin1-decoded source text is `ONE_BYTE`-encoded, slices of it are too. In comparison, slices of `sourceText` may be `ONE_BYTE` or `TWO_BYTE`. If a file's source is pure ASCII, it'll be `ONE_BYTE`, if source contains any non-ASCII characters, it'll be `TWO_BYTE`. Files in a repo will likely be a mix of both, which makes strings returned from `deserializeStr` and placed in the AST a mix too. This in turn makes functions (e.g. lint rule visitors) polymorphic. V8 cannot optimize them as aggressively as if they see only `ONE_BYTE` strings.

We cannot make sure that all strings returned by `deserializeStr` are `ONE_BYTE`. Some string may contain non-ASCII characters, and they *have* to be represented in `TWO_BYTE` form. But we can minimize it - now only strings which *themselves* contain non-ASCII characters are `TWO_BYTE`, whereas before they would be if the source text as a whole contains a single non-ASCII byte.

Code which accesses `Identifier` names, for example will exclusively see `ONE_BYTE` strings and will be more heavily optimized, because Unicode `Identifier`s are rarer than hen's teeth in real-world code.

### Remove string-concatenation loop

Previously strings which are outside of source text were assembled byte-by-byte in a loop via concatenation.

Instead, check that all the bytes are ASCII first, copy them into an array and pass that array to `String.fromCharCode` with `fromCharCode.apply(null, array)`.

To avoid allocating a fresh array every time, hold a stock of arrays for all string lengths that this path can require, and reuse them.

This is a variation on the approach that #20883 took, but without the massive switch. This produces much tighter assembly, and avoids regressing the fast path due making `deserializeStr` a very large function.

Despite the complexity, and multiple operations, [this is up to 3x faster](https://github.com/overlookmotel/oxc-raw-str-bench/blob/4f96275efa9a35d5d27615abb27f21a137149cc0/README.md#apply30-vs-switch30) than the switch approach, and gives an average 30% speed-up.

### Increase native call threshold

The above optimizations make the slow path much faster. This shifts the tipping point at which it's faster to make a native call to `TextDecoder.decode` from 9 bytes to 64 bytes. Most strings now avoid the native call and stay in JS code which is heavily optimized by Turbofan.

The tipping point of 64 is something of a guesstimate. Benchmarking shows its in the right ballpark, but we could finesse it, and probably squeeze out another couple of %.

## Credit

The Latin1 string technique was cooked up by @joshuaisaact in https://github.com/overlookmotel/oxc-raw-str-bench/pull/1. All credit to him for this masterstroke which cracks the whole problem!
2026-04-03 22:05:21 +00:00
overlookmotel 55e1e9b1d2 perf(napi/parser, linter/plugins): initialize vars as 0 (#21020)
Tiny optimization to raw transfer string deserialization. Initialize vars which contain integers as `0`. This may help V8 optimize usage of these variables as they always have SMI type.

The effect, if there is one, is too small to register on benchamarks. But it's certainly not a regression, and may be a tiny gain, so why not?
2026-04-03 22:05:21 +00:00
overlookmotel c25ef02453 perf(napi/parser, linter/plugins): simplify branch condition in deserializeStr (#21019)
Follow-on after #20834.

Simplify the branch condition in `deserializeStr` for detemining if can take the fast path of just slicing `sourceText`. There's no need to check `sourceIsAscii`, just compare the offset to `firstNonAsciiPos` (the position in buffer of first non-ASCII byte in source code). When source is 100% ASCII, `firstNonAsciiPos = sourceEndPos`, so `pos < firstNonAsciiPos` passes for all positions in source.

The implementation is different for parser and for Oxlint, as the source text sits in a different location in buffer - at the start in parser, at the end in Oxlint - but the principle is the same in both.

[Benchmarking](https://github.com/overlookmotel/oxc-raw-str-bench) showed this speeds up `deserializeStr` by a small percentage.
2026-04-03 22:05:20 +00:00
overlookmotel 9f494c3bc8 perf(napi/parser, linter/plugins): raw transfer use String.fromCharCode in string decoding (#21018)
Small optimization to raw transfer string deserialization. Use `String.fromCharCode`instead of `String.fromCodePoint`. [Benchmarking](https://github.com/overlookmotel/oxc-raw-str-bench) showed it's slightly faster, as it's simpler - doesn't need to handle astral code points.
2026-04-03 22:05:20 +00:00
overlookmotel 15546c0261 refactor(napi/parser, linter/plugins): shorten raw transfer deserializers for Options (#20924)
Refactor. Just shorten the code generated for deserializing `Option`s in raw transfer deserializer.
2026-04-01 12:08:08 +00:00
overlookmotel 0503a78b8b perf(napi/parser, linter/plugins): faster deserialization of raw fields (#20923)
`raw` field of `NumericLiteral`, `StringLiteral`, `BigIntLiteral`, `RegExpLiteral`, and `JSXText` are slices of source text.

So in raw transfer deserializer, skip going through `deserializeStr`, which can be slow when source contains any non-ASCII characters. Instead just slice the string from the source text directly with `sourceText.slice(start, end)`.

String decoding is the slowest part of raw transfer, so this should be a significant speed gain.
2026-04-01 12:08:08 +00:00
Joshua Tuddenham a24f75e9e8 perf(napi/parser): optimize string deserialization for non-ASCII sources (#20834)
**AI Disclosure:** Developed with Claude Code (Opus). The winning
approach came out of an automated experiment loop
([auto-claude](https://github.com/joshuaisaact/auto-claude)) — I was
looking for a tight feedback loop to test the tool on and the `TODO:
Find best switch-over point` comment in `deserializeStr` caught my eye.
20 experiments, keep-or-revert on each (all 20 summarized in an
expandable section at the bottom). All code reviewed and understood.
Happy to close this if it's not useful or doesn't meet the bar.

## Why

I was profiling the raw transfer deserialization path (`node --prof` on
`checker.ts`) and noticed `StringAdd_CheckNone` at 13.7% of time — the
single hottest function. It comes from the byte-by-byte `out +=
fromCodePoint(c)` loop in `deserializeStr` when `sourceIsAscii` is
false.

The thing is, `sourceIsAscii` is false for almost everything. All 5 NAPI
bench fixtures are non-ASCII. `checker.ts` has literally one Bengali
character at position 2.1M out of 2.9M. That one character disables the
fast `substr` path for all ~148K strings.

## What

Two changes to the generator
(`tasks/ast_tools/src/generators/raw_transfer.rs`) — that's the only
file with real changes. The 9 generated JS files in the diff are the
mechanical output of `cargo run -p oxc_ast_tools`.

**1. `firstNonAsciiPos` scan at init** — On non-ASCII sources, find the
first non-ASCII byte once upfront. Strings ending before that position
can still use `sourceText.substr()` since byte offsets equal char
offsets in the ASCII prefix. For `checker.ts` this covers 73% of the
file, for `pdf.mjs` 98%.

**2. Lower TextDecoder threshold from 50 to 9** — The existing TODO
asked for the right switch-over point. Experimentally, 9 is the sweet
spot: `TextDecoder` beats the `fromCodePoint` concat loop for strings of
10+ bytes, and the concat loop is still faster for very short strings
where the native call overhead dominates.

Benchmarked on the `complicated()` test set (5 rounds of 30 iters,
dropping round 1 for JIT warmup):

```
              Before    After
checker.ts    26.7ms    13.0ms   -51%
cal.com.tsx   15.4ms     9.1ms   -41%
antd.js       53.8ms    44.7ms   -17%
pdf.mjs        3.8ms     4.3ms   noise
binder.ts      0.7ms     0.5ms   noise
Total        100.5ms    71.6ms   -29%
```

Also verified across 15 files — the 10 non-ASCII files above plus 5
ASCII files (`react.development.js`, `binder.ts`, `moment.js`,
`jquery.js`, `vue.js`). ASCII files are unchanged (our code only touches
the non-ASCII path). -16% across non-ASCII files, no regressions.

## References

- Addresses the `TODO: Find best switch-over point` in
`STR_DESERIALIZER_BODY`
- Related to perf goals in #19918

<details>
<summary>I appreciate this is already a bloated PR description (sorry
@overlookmotel) but given the fairly unusual approach I thought you
might want to see a very short summary of all 20 experiments Claude
clauded through:</summary>

The loop works like this: edit code, benchmark, keep if faster, `git
reset --hard` if not. Metric is total deserialization time across the
benchmark corpus.

**Baseline: 97.7ms** (5-file corpus, all non-ASCII)

| # | Idea | Result | Verdict |
|---|------|--------|---------|
| 1 | Always use TextDecoder, delete the loop entirely | 99.2ms |
Revert. TextDecoder's ~78ns fixed overhead kills short strings. |
| 2 | Lower TextDecoder threshold from 50 to 10 (ts.js only) | 89.8ms |
**Keep.** First real win — moves 50% of strings off the concat loop. |
| 3 | Threshold 5 | 91.3ms | Revert. Too aggressive, too many short
strings go to TextDecoder. |
| 4 | Threshold 8 | 95.0ms | Revert. Worse than 10. |
| 5 | Threshold 12 | 91.2ms | Revert. Worse than 10. |
| 6 | Threshold 10 + unrolled `switch` on len for inline
`String.fromCharCode(uint8[pos], ...)` | 90.9ms | Revert. Switch
dispatch overhead eats the gain. |
| 7 | Threshold 10 + special fast path for len=1 | 94.8ms | Revert.
Extra branch hurts more than the 1-byte optimization helps. |
| 8 | Threshold 10 + accumulate char codes in array, single
`fromCharCode.apply` at end | 90.5ms | Revert. Array allocation
overhead. |
| 9 | Threshold 15 | 93.2ms | Revert. 10 is still the sweet spot. |
| 10 | Unrolled ASCII check for bytes 1-4, TextDecoder for 5+ | 94.7ms |
Revert. Branching overhead. |
| 11 | **Apply threshold 10 to js.js too** (had only been changing
ts.js) | 82.5ms | **Keep.** Facepalm moment — antd.js uses the JS
deserializer. |
| 12 | Threshold 9 in both files | 77.9ms | **Keep.** New best. |
| 13 | Threshold 7 | 79.0ms | Revert. 9 wins. |
| 14 | Threshold 11 | 84.8ms | Revert. 9 confirmed. |
| 15 | Replace `fromCodePoint` with `String.fromCharCode` in the loop |
82.5ms | Revert. V8 optimizes the pre-extracted `fromCodePoint` better.
|
| 16 | Various unrolled `fromCharCode` approaches for short strings | —
| Abandoned, too complex for marginal gain. |
| 17 | Always TextDecoder for non-source strings (remove loop) | 96.1ms
| Revert. Confirms the short-string loop IS valuable for 1-9 bytes. |
| 18 | `Buffer.from().toString()` instead of TextDecoder | 80.4ms |
Revert. TextDecoder is faster. |
| 19 | `firstNonAsciiPos` only (use substr before it, TextDecoder after,
no loop) | 84.9ms | Revert. checker.ts loved it (-34%) but antd.js hated
it (+20%) because its first non-ASCII byte is at 1.3%. |
| 20 | **`firstNonAsciiPos` + threshold 9 + keep the loop** | 76.4ms |
**Keep.** Best of both worlds — substr where possible, TextDecoder for
medium strings, loop for short. |

Three things I (Claude) learned:
- Experiment 11 was the biggest single win and it was just... applying
the change to the other file. Embarrassing.
- Every attempt to replace the `fromCodePoint` loop for short strings
(1-9 bytes) made things worse. The loop is genuinely good for that
range.
- `firstNonAsciiPos` only works when combined with the threshold change.
On its own it hurts files where non-ASCII appears early (antd.js,
cal.com.tsx).

</details>
2026-03-30 16:19:17 +01:00
overlookmotel 9a622c79b4 perf(linter/plugins): lazy deserialize tokens and comments (#20474)
Performance improvement to tokens and comments APIs.

## The problem

Previously, all tokens and comments methods would deserialize *all* tokens/comments into an array of `Token` / `Comment` / `Token | Comment` objects, and then binary search through those arrays to find the token(s) / comment(s) they're looking for.

This has 2 major disadvantages:

1. Files typically contain *a lot* of tokens (even more than the number of AST nodes). Deserializing them all is very costly (up to 30% of total Oxlint runtime when run with only a JS rule which just calls a tokens-related method).

2. The binary searches these methods do are quite expensive. Even in TurboFan-optimized code, accessing `token.start` involves getting pointer to the `Token` object from the `tokens` array, an "is this object a `Token`?" safety check, then reading the `start` field from the `Token` - all just to access a single `u32`, and that happens over and over.

## This PR's solution

Solve both these problems by making tokens and comments methods read `start` / `end` offsets directly from the buffers which contain the tokens/comments data.

This data is tightly packed in memory, and strongly typed (read from `Uint32Array`s), so getting `start` / `end` of a token requires no indirection and no type checks.

More importantly, it removes the need to deserialize all tokens / comments upfront. The desired token(s) are located, touching only the buffer, and then *only* the ones which need to be returned to rule code are deserialized into JS objects.

If a rule accesses `ast.tokens`, `ast.comments`, or `sourceCode.tokensAndComments` then all tokens / comments need to be deserialized, as they're all returned to the rule as an array - but that's unavoidable. This PR doesn't make that any cheaper, but it doesn't make it measurably more costly either.

But where no rule requires the full array of tokens / comments, and they only use token/comment search methods (e.g. `getFirstToken`, `getCommentsBefore`), a great deal of work will be saved. This covers the vast majority of rules.

## Implementation details

The main complication is the `includeComments` option to tokens methods. When `true`, search needs to be over a combined set of both tokens and comments.

When `includeComments: true` option is passed to a tokens method, a buffer is created containing data about all tokens and comments, interleaved in source code order. This buffer can then be used for binary search in tokens methods.

Whether each token / comment has been deserialized already or not is tracked by a "deserialized" flag in the tokens/comments buffers. Each token / comment in the buffer is 16 bytes. This flag lives in byte 15. For tokens, this byte is always already 0 in the buffer when it arrives from Rust side. For comments, we manually set `comment.content = CommentContent::None;` for every comment on Rust side. `comment.content` is positioned at byte 15 in the `Comment` struct, and `CommentContent::None` is stored as 0.

## Possible future improvements

### SoA storage

Binary search operates only on `start` field of tokens / comments, which are 16 bytes apart in the buffer. It would be more efficient if tokens were stored in struct-of-arrays (SoA) style so all `start` values were tightly packed together. This would reduce CPU cache misses in the hot loops of binary searches.

### Pre-compute tokens-and-comments buffer on Rust side

The buffer containing tokens and comments, required to support `includeComments: true`, is currently generated on JS side (but lazily). We could move that to Rust side, which would be faster. However, it might be redundant work in many cases because the buffer is only required if a rule uses `includeComments: true`.

We could alternatively keep the laziness optimization, by calling back into Rust to build the buffer on demand - but JS-Rust calls have a cost too. Maybe communicating via `Atomics` would be faster than an actual function call?

If we had a way to share buffers with WASM, optimal solution might be to generate the buffer lazily (as now) but in WASM, which would be faster for this kind of pure number-crunching, but without the overhead of calling into Rust.
2026-03-21 12:24:11 +00:00
overlookmotel c6ea0a068d perf(ast): place NodeId field after Span in structs (#20584)
Alter the code in `ast_tools` which re-orders struct fields, to make `NodeId` always occupy a consistent "slot" in every AST node - at byte position 8, just after `Span`.

This will be a sizeable gain once we start utilizing the `NodeId` fields more. Because, for example, every variant of `Expression` has its `NodeId` stored in same location as all the other variants, `Expression::node_id()` is just 1 operation - a pointer read - rather than a nest of branches, or a lookup table.

Also alter the algorithm for ordering struct fields to fill in the 4-byte gap after `NodeId` field with other field(s), to avoid excess padding.

The new algorithm also prioritizes keeping fields in definition order as much as possible, rather than sorting them strictly in order of size and alignment. This is mildly advantageous because field definition order is the order the AST is walked in, so it avoids bouncing between cache lines while iterating through the fields of a struct when visiting the AST.

No types change size in the process. Fields remain packed to keep type sizes the minimum they can be - they just change order.
2026-03-20 22:15:56 +00:00
overlookmotel 1be1ebe1ae refactor(ast_tools): search all crates which depend on oxc_ast_macros (#20577)
One annoyance with `ast_tools`, our codegen, has been that when you want it to generate code from types in some file, you had to manually add the file paths to a list in `ast_tools` itself.

This PR removes that list and instead finds crates in the monorepo which depend on `oxc_ast_macros` crate, via `cargo metadata`, and searches all files in those crates for types with `#[ast]` attributes etc - types that it need to act on.

There is no longer a list to keep updated, you just mark types `#[ast]` and the codegen will find them.

This produces some churn in generated files. None of the generated code changes, it just gets re-ordered in some places, due to the different order that files get found in now. This order is deterministic, and it should not change again.
2026-03-20 22:15:54 +00:00
overlookmotel d176eccdbf perf(napi/parser, oxlint/plugins): shorten deserializer for WithClause (#20575)
Small optimization to ESTree serialization and raw transfer. Shave off a little work when serializing/deserializing `WithClause`.
2026-03-20 20:40:34 +00:00
overlookmotel 33993feb98 refactor(estree): simplify ESTree converter for VariableDeclarator (#20574)
Belated follow-on after #15925.

`VariableDeclarator` doesn't need a whole-struct ESTree converter. Switch to a converter just for the `id` field, to reduce custom conversion code.
2026-03-20 20:05:10 +00:00
overlookmotel 9cd612f1e9 perf(linter/plugins): recycle comment objects (#20362)
Apply the same optimization as #19978 to comments - hold a pool of `Comment` objects, and re-use those objects rather than creating new objects each time.

Same as with `Token`s, `loc` property is a getter which calculates `loc` lazily, and caches it in a private property.
2026-03-14 12:19:27 +00:00
overlookmotel e4aa5b51c1 docs(parser/napi, linter/plugins): add JSDoc comments to raw transfer constants (#20286)
Comments-only change. Add JSDoc comments to the constants related to raw transfer in generated files `constants.js` / `constants.ts`. There are now a lot of constants, and it wasn't clear what they all mean.
2026-03-12 12:46:02 +00:00
overlookmotel 92cfb14aa4 fix(linter/plugins): fix types for walkProgram and walkProgramWithCfg (#20081)
Fix a mistake in our (internal) types: `walkProgramWithCfg` was missing that elements can be `CfgVisitFn`.

Also refactor the types for `walkProgram`, importing the original `VisitFn` and `EnterExit` types from where they're defined.
2026-03-06 18:04:24 +00:00