Large change to how raw transfer stores metadata about the allocations it uses.
Previously the two metadata structures `RawTransferMetadata` and `FixedSizeAllocator` after the chunk's `ChunkFooter`.
This had a few problems:
1. It was confusing and unwieldy - a recipe for bugs.
2. `ChunkFooter`'s memory was exposed on JS side - a bit unsafe as altering bytes in this region could easily trigger UB.
3. It entwined `Arena` (which is just the thing that allocates) with the details of exactly _what_ raw transfer allocates.
4. Imposed annoying alignment requirements, because `ChunkFooter` must be aligned on 16, and so anything after it must ensure it doesn't break that invariant.
Old layout:
```
WHOLE BLOCK - aligned on 4 GiB
<-----------------------------------------------------> Allocated block (`BLOCK_SIZE` bytes)
ALLOCATOR
<-----------------------------------------> `Allocator` chunk (`CHUNK_SIZE` bytes)
<----> `ChunkFooter` (aligned on 16)
<-----------------------------------> `Allocator` chunk data storage (for AST)
(`ACTIVE_SIZE` bytes)
METADATA
<----> `RawTransferMetadata`
<----> `FixedSizeAllocatorMetadata`
BUFFER SENT TO JS
<-----------------------------------------------> Buffer sent to JS (`BUFFER_SIZE` bytes)
```
This PR moves `RawTransferMetadata` and `FixedSizeAllocatorMetadata` into the chunk itself. New layout:
```
WHOLE BLOCK - size 2 GiB - 16, aligned on 4 GiB
<-----------------------------------------------------> Allocated block (`BLOCK_SIZE` bytes)
ARENA
<-----------------------------------------------------> Chunk (fills whole block)
<--------------------------------------> Allocatable region for AST (`ACTIVE_SIZE` bytes)
<---> `RawTransferMetadata`
<---> `FixedSizeAllocatorMetadata`
<---> `ChunkFooter` (aligned on 16, last in block)
BUFFER SENT TO JS
<-------------------------------------------> Buffer sent to JS (`BUFFER_SIZE` bytes)
```
`FixedSizeAllocatorMetadata` and `ChunkFooter` are no longer in the region which is shared with JS side. As far as `Arena` is concerned, they're now just some data (like any other data) which is allocated in the arena.
Also:
- Introduce more consistency to the naming of constants which specify the size and position of these various data structures in the arena.
- Add more const assertions to ensure everything is laid out and aligned as it should be.
Use `int32` (`Int32Array`) instead of `uint32`(`Uint32Array`) for getting data from buffer in raw transfer deserializer. The buffer is 2 GiB in size, so all offsets within the buffer can be represented as a `u31` which can be stored in an `i32` without any loss, and with no risk of being interpreted as negative numbers.
Same as in #21129, the advantage of fetching offsets from an `Int32Array` is that V8 statically knows the value can be stored in an SMI, its native integer type, without any range checks.
Performance improvement to tokens and comments APIs.
## The problem
Previously, all tokens and comments methods would deserialize *all* tokens/comments into an array of `Token` / `Comment` / `Token | Comment` objects, and then binary search through those arrays to find the token(s) / comment(s) they're looking for.
This has 2 major disadvantages:
1. Files typically contain *a lot* of tokens (even more than the number of AST nodes). Deserializing them all is very costly (up to 30% of total Oxlint runtime when run with only a JS rule which just calls a tokens-related method).
2. The binary searches these methods do are quite expensive. Even in TurboFan-optimized code, accessing `token.start` involves getting pointer to the `Token` object from the `tokens` array, an "is this object a `Token`?" safety check, then reading the `start` field from the `Token` - all just to access a single `u32`, and that happens over and over.
## This PR's solution
Solve both these problems by making tokens and comments methods read `start` / `end` offsets directly from the buffers which contain the tokens/comments data.
This data is tightly packed in memory, and strongly typed (read from `Uint32Array`s), so getting `start` / `end` of a token requires no indirection and no type checks.
More importantly, it removes the need to deserialize all tokens / comments upfront. The desired token(s) are located, touching only the buffer, and then *only* the ones which need to be returned to rule code are deserialized into JS objects.
If a rule accesses `ast.tokens`, `ast.comments`, or `sourceCode.tokensAndComments` then all tokens / comments need to be deserialized, as they're all returned to the rule as an array - but that's unavoidable. This PR doesn't make that any cheaper, but it doesn't make it measurably more costly either.
But where no rule requires the full array of tokens / comments, and they only use token/comment search methods (e.g. `getFirstToken`, `getCommentsBefore`), a great deal of work will be saved. This covers the vast majority of rules.
## Implementation details
The main complication is the `includeComments` option to tokens methods. When `true`, search needs to be over a combined set of both tokens and comments.
When `includeComments: true` option is passed to a tokens method, a buffer is created containing data about all tokens and comments, interleaved in source code order. This buffer can then be used for binary search in tokens methods.
Whether each token / comment has been deserialized already or not is tracked by a "deserialized" flag in the tokens/comments buffers. Each token / comment in the buffer is 16 bytes. This flag lives in byte 15. For tokens, this byte is always already 0 in the buffer when it arrives from Rust side. For comments, we manually set `comment.content = CommentContent::None;` for every comment on Rust side. `comment.content` is positioned at byte 15 in the `Comment` struct, and `CommentContent::None` is stored as 0.
## Possible future improvements
### SoA storage
Binary search operates only on `start` field of tokens / comments, which are 16 bytes apart in the buffer. It would be more efficient if tokens were stored in struct-of-arrays (SoA) style so all `start` values were tightly packed together. This would reduce CPU cache misses in the hot loops of binary searches.
### Pre-compute tokens-and-comments buffer on Rust side
The buffer containing tokens and comments, required to support `includeComments: true`, is currently generated on JS side (but lazily). We could move that to Rust side, which would be faster. However, it might be redundant work in many cases because the buffer is only required if a rule uses `includeComments: true`.
We could alternatively keep the laziness optimization, by calling back into Rust to build the buffer on demand - but JS-Rust calls have a cost too. Maybe communicating via `Atomics` would be faster than an actual function call?
If we had a way to share buffers with WASM, optimal solution might be to generate the buffer lazily (as now) but in WASM, which would be faster for this kind of pure number-crunching, but without the overhead of calling into Rust.
Alter the code in `ast_tools` which re-orders struct fields, to make `NodeId` always occupy a consistent "slot" in every AST node - at byte position 8, just after `Span`.
This will be a sizeable gain once we start utilizing the `NodeId` fields more. Because, for example, every variant of `Expression` has its `NodeId` stored in same location as all the other variants, `Expression::node_id()` is just 1 operation - a pointer read - rather than a nest of branches, or a lookup table.
Also alter the algorithm for ordering struct fields to fill in the 4-byte gap after `NodeId` field with other field(s), to avoid excess padding.
The new algorithm also prioritizes keeping fields in definition order as much as possible, rather than sorting them strictly in order of size and alignment. This is mildly advantageous because field definition order is the order the AST is walked in, so it avoids bouncing between cache lines while iterating through the fields of a struct when visiting the AST.
No types change size in the process. Fields remain packed to keep type sizes the minimum they can be - they just change order.
Apply the same optimization as #19978 to comments - hold a pool of `Comment` objects, and re-use those objects rather than creating new objects each time.
Same as with `Token`s, `loc` property is a getter which calculates `loc` lazily, and caches it in a private property.
Comments-only change. Add JSDoc comments to the constants related to raw transfer in generated files `constants.js` / `constants.ts`. There are now a lot of constants, and it wasn't clear what they all mean.
Builds on #19497. Use the tokens generated by Oxc's parser in linter plugins, instead of running TS parser to tokenize source.
This is a sizeable perf gain, and also allows us to remove TS parser from `oxlint`'s bundle (#19531).
Additionally, our parser is more accurate than TypeScript - we handle HTML comments and space before slash in closing JSX elements (`< /div>`) - so this fixes a few conformance tests.
Previously we had to store source text at start of buffers sent to JS via raw transfer.
#18376 made changes to how raw transfer deserializer handles strings, in order to support files containing a BOM. Building on that, we're now able to remove the requirement that source text be at start of the buffer entirely.
This PR changes the deserializer used in Oxlint JS plugins to accept source text being anywhere in the buffer, *as long as no other strings are after it*. In practice this just means that the source text must be allocated before anything else, which is easy to satisfy.
Now the source text can be allocated with just the usual safe `allocator.alloc_str(source_text)` method.
This change removes a ton of dodgy workarounds and unsafe code we used previously to get source text at the start of buffer. It makes the code less labyrinthine and far less likely a slip up can inadvertently introduce UB.
Note: In `napi/parser`, source text still *is* at start of the buffer, as that's simpler and more efficient when the source text is written into the buffer on JS side. This change only affects Oxlint.
Closes#12526.
Handle BOM on start of files in the same way that ESLint does - do not include it in the source text on JS side, but `context.sourceCode.hasBOM` evaluates to `true`.
Method:
* Alter `program.source_text` to trim off the BOM before passing AST to JS side.
* Add a `has_bom` flag to `RawTransferMetadata`.
* Add ability to add an offset in the conversion from UTF-8 to UTF-16 spans.
The result is that the file as it's seen on JS side is as if the BOM didn't exist (except for the `hasBOM` flag). Spans are converted accordingly in JS-side AST, and converted back when passing diagnostics back to Rust.
Previous raw transfer required that the source text start exactly at the start of the buffer. In linter, relax this restriction and handle when source text is stored elsewhere in the buffer.
This is necessary for stripping BOM from source (#18376), as then the source text used on JS side starts 3 bytes after the start of the buffer.
When parsing tokens for JS plugins, the TypeScript parser was always using `ScriptKind.TSX`, regardless of file extension. This caused TypeScript to incorrectly parse generic arrow functions like `<T>() => {}` in `.ts` files, creating bogus `JsxText` tokens that overlap with comments.
This PR adds an `is_jsx` flag to `RawTransferMetadata`. Rust side sets the flag depending on source type, and JS side uses that information to pass the correct `ScriptKind` for the file to TypeScript parser.
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: overlookmotel <theoverlookmotel@gmail.com>