Repository navigation
fix(deps): update dependency @huggingface/transformers to v4 - #553
Open
renovate[bot] wants to merge 1 commit into
Open
renovate[bot] wants to merge 1 commit into
renovate[bot] wants to merge 1 commit into
Conversation
renovate
Bot
force-pushed
the
renovate/huggingface-transformers-4.x
branch
from
August 12, 2026 03:33
459cc2c to
befa76c
Compare
sharevb
force-pushed
the
chore/all-my-stuffs
branch
2 times, most recently
from
September 6, 2026 07:59
cc0812b to
62ab0bf
Compare
sharevb
force-pushed
the
chore/all-my-stuffs
branch
2 times, most recently
from
September 27, 2026 20:37
3ae31a6 to
90b7793
Compare
renovate
Bot
force-pushed
the
renovate/huggingface-transformers-4.x
branch
from
October 3, 2026 23:16
befa76c to
f6ff361
Compare
renovate
Bot
force-pushed
the
renovate/huggingface-transformers-4.x
branch
from
October 7, 2026 00:57
f6ff361 to
e8c35fc
Compare
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR contains the following updates:
3.7.6→4.3.1Release Notes
huggingface/transformers.js (@huggingface/transformers)
v4.3.1Compare Source
What's new?
This release adds support for EmbeddingGemma 2. EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.
Text search
The
feature-extractionpipeline returns normalized embeddings, so the dot product of two embeddings is their cosine similarity:Images, audio and video
For other modalities, use the processor and the model directly. Every input maps into the same embedding space, so any embedding can be compared with any other: here, text queries against an image, an audio clip and a video.
To embed several items of a modality at once, pass one list per input:
processor(null, [[image1], [image2]])returns two image embeddings, while a flat list of images,processor(null, [image1, image2]), is a single input made of both images.Full Changelog: huggingface/transformers.js@4.3.0...4.3.1
v4.3.0Compare Source
🚀 Transformers.js v4.3 — Structured Output, New Models, WebGPU upgrade, Documentation Overhaul
This release adds structured output, three new model architectures, WebGPU support for Safari 26+, and a documentation overhaul. We also upgraded ONNX Runtime to the latest version.
What's new?
Structured output
Constrain generation to a JSON schema, JSON object, or regular expression with the experimental, dependency-free
@huggingface/transformers-structured-outputpackage in #1758.For example, classify customer feedback into a fixed set of sentiments and topics:
Currently supports one generated sequence at a time. Set a sufficient token budget so the output can finish.
New models
Browser and storage improvements
requestFileHandle()and allow all origins by @tomayac in #1709 and #1716import.metawarnings by @nico-martin in #1760Fixes
progress_callbackis active by @anishesg in #1664num_logits_to_keepis always set to1for generation in #1681RawAudio.toBlob()to respect typed array byte offsets and lengths by @yushuosun in #1712Documentation and maintenance
New contributors
Thanks to our first-time contributors:
Full changelog: huggingface/transformers.js@4.2.0...4.3.0
v4.2.0Compare Source
🚀 Transformers.js v4.2 — Tool calling, simpler internals, and privacy filtering
toolstoTextGenerationPipelinein #1655inputMetadataAPI for simplified internals in #1657Full Changelog: 4.1.0...4.2.0
v4.1.0Compare Source
🚀 Transformers.js v4.1 — Gemma 4, KV cache improvements, and new quantization dtypes
past_key_valuesvia pipeline function) in #1638q1,q1f16,q2, andq2f16data types in #1647Full Changelog: 4.0.0...4.1.0
v4.0.1Compare Source
What's new?
Full Changelog: 4.0.0...4.0.1
v4.0.0Compare Source
🚀 Transformers.js v4
We're excited to announce that Transformers.js v4 is now available on NPM! After a year of development (we started in March 2025 🤯), we're finally ready for you to use it.
Links: YouTube Video, Blog Post, Demo Collection
New WebGPU backend
The biggest change is undoubtedly the adoption of a new WebGPU Runtime, completely rewritten in C++. We've worked closely with the ONNX Runtime team to thoroughly test this runtime across our ~200 supported model architectures, as well as many new v4-exclusive architectures.
In addition to better operator support (for performance, accuracy, and coverage), this new WebGPU runtime allows the same transformers.js code to be used across a wide variety of JavaScript environments, including browsers, server-side runtimes, and desktop applications. That's right, you can now run WebGPU-accelerated models directly in Node, Bun, and Deno!
We've proven that it's possible to run state-of-the-art AI models 100% locally in the browser, and now we're focused on performance: making these models run as fast as possible, even in resource-constrained environments. This required completely rethinking our export strategy, especially for large language models. We achieve this by re-implementing new models operation by operation, leveraging specialized ONNX Runtime Contrib Operators like com.microsoft.GroupQueryAttention, com.microsoft.MatMulNBits, or com.microsoft.QMoE to maximize performance.
For example, adopting the com.microsoft.MultiHeadAttention operator, we were able to achieve a ~4x speedup for BERT-based embedding models.
New models
Thanks to our new export strategy and ONNX Runtime's expanding support for custom operators, we've been able to add many new models and architectures to Transformers.js v4. These include popular models like GPT-OSS, Chatterbox, GraniteMoeHybrid, LFM2-MoE, HunYuanDenseV1, Apertus, Olmo3, FalconH1, and Youtu-LLM. Many of these required us to implement support for advanced architectural patterns, including Mamba (state-space models), Multi-head Latent Attention (MLA), and Mixture of Experts (MoE). Perhaps most importantly, these models are all compatible with WebGPU, allowing users to run them directly in the browser or server-side JavaScript environments with hardware acceleration. We've released several Transformers.js v4 demos so far... and we'll continue to release more!
Additionally, we've added support for larger models exceeding 8B parameters. In our tests, we've been able to run GPT-OSS 20B (q4f16) at ~60 tokens per second on an M4 Pro Max.
New features
ModelRegistry
The new
ModelRegistryAPI is designed for production workflows. It provides explicit visibility into pipeline assets before loading anything: list required files withget_pipeline_files, inspect per-file metadata withget_file_metadata(quite useful to calculate total download size), check cache status withis_pipeline_cached, and clear cached artifacts withclear_pipeline_cache. You can also query available precision types for a model withget_available_dtypes. Based on this new API,progress_callbacknow includes aprogress_totalevent, making it easy to render end-to-end loading progress without manually aggregating per-file updates.See `ModelRegistry` examples
New Environment Settings
We also added new environment controls for model loading.
env.useWasmCacheenables caching of WASM runtime files (when cache storage is available), allowing applications to work fully offline after the initial load.env.fetchlets you provide a custom fetch implementation for use cases such as authenticated model access, custom headers, and abortable requests.See env examples
Improved Logging Controls
Finally, logging is easier to manage in real-world deployments. ONNX Runtime WebGPU warnings are now hidden by default, and you can set explicit verbosity levels for both Transformers.js and ONNX Runtime. This update, also driven by community feedback, keeps console output focused on actionable signals rather than low-value noise.
See `logLevel` example
is_cached/is_pipeline_cached, closes #1554 by @nico-martin in #1559progress_totalevents fromPreTrainedModel.from_pretrained()by @xenova in #1615Repository Restructuring
Developing a new major version gave us the opportunity to invest in the codebase and tackle long-overdue refactoring efforts.
PNPM Workspaces
Until now, the GitHub repository served as our npm package. This worked well as long as the repository only exposed a single library. However, looking to the future, we saw the need for various sub-packages that depend heavily on the Transformers.js core while addressing different use cases, like library-specific implementations, or smaller utilities that most users don't need but are essential for some.
That's why we converted the repository to a monorepo using pnpm workspaces. This allows us to ship smaller packages that depend on
@huggingface/transformerswithout the overhead of maintaining separate repositories.Modular Class Structure
Another major refactoring effort targeted the ever-growing models.js file. In v3, all available models were defined in a single file spanning over 8,000 lines, becoming increasingly difficult to maintain. For v4, we split this into smaller, focused modules with a clear distinction between utility functions, core logic, and model-specific implementations. This new structure improves readability and makes it much easier to add new models. Developers can now focus on model-specific logic without navigating through thousands of lines of unrelated code.
Examples Repository
In v3, many Transformers.js example projects lived directly in the main repository. For v4, we've moved them to a dedicated repository, allowing us to maintain a cleaner codebase focused on the core library. This also makes it easier for users to find and contribute to examples without sifting through the main repository.
Prettier
We updated the Prettier configuration and reformatted all files in the repository. This ensures consistent formatting throughout the codebase, with all future PRs automatically following the same style. No more debates about formatting... Prettier handles it all, keeping the code clean and readable for everyone.
Standalone Tokenizers.js Library
A frequent request from users was to extract the tokenization logic into a separate library, and with v4, that's exactly what we've done. @huggingface/tokenizers is a complete refactor of the tokenization logic, designed to work seamlessly across browsers and server-side runtimes. At just 8.8kB (gzipped) with zero dependencies, it's incredibly lightweight while remaining fully type-safe.
See example code
This separation keeps the core of Transformers.js focused and lean while offering a versatile, standalone tool that any WebML project can use independently.
New build system
We've migrated our build system from Webpack to esbuild, and the results have been incredible. Build times dropped from 2 seconds to just 200 milliseconds, a 10x improvement that makes development iteration significantly faster. Speed isn't the only benefit, though: bundle sizes also decreased by an average of 10% across all builds. The most notable improvement is in transformers.web.js, our default export, which is now 53% smaller, meaning faster downloads and quicker startup times for users.
Improved types
We've made several quality-of-life improvements across the library. The type system has been enhanced with dynamic pipeline types that adapt based on inputs, providing better developer experience and type safety.
Bug fixes
stopping_criteriamissing from generation pipelines by @xenova in #1523content-lengthheader in COSResponseby @tomayac in #1572Documentation improvements
Miscellaneous improvements
skip_special_tokens: falseby @xenova in #1520New Contributors
Full Changelog: huggingface/transformers.js@3.8.1...4.0.0
v3.8.1Compare Source
What's new?
Full Changelog: huggingface/transformers.js@3.8.0...3.8.1
v3.8.0Compare Source
🚀 Transformers.js v3.8 — SAM2, SAM3, EdgeTAM, Supertonic TTS
Add support for EdgeTAM in #1454
Add support for Supertonic TTS in #1459
Example:
Add support for SAM2 and SAM3 (Tracker) in #1461
Remove Metaspace add_prefix_space logic in #1451
ImageProcessor preprocess uses image_std for fill value by @NathanKolbas in #1455
New Contributors
Full Changelog: huggingface/transformers.js@3.7.6...3.8.0
Configuration
📅 Schedule: (in timezone Europe/Paris)
* 0-3 * * *)* 0-3 * * *)🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.
♻ Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.
🔕 Ignore: Close this PR and you won't be reminded about this update again.
This PR was generated by Mend Renovate. View the repository job log.