Best browser extensions for local LLMs in 2026
If you want to chat with a local model in a sidebar, Page Assist is the most complete free option: open source, works with Ollama and LM Studio, and needs no extra Ollama setup on localhost. If you want the local model to write into the text field you are using, such as an email reply or a form, ShroomPen and Scramble do that. NativeMind, SurfMind and Lumos are sidebar assistants with different trade-offs. All of them are free.
Published
The comparison
The biggest difference between these extensions is where the answer ends up. Most are chats: you ask in a sidebar and copy the answer out. A few put the answer into the page for you.
| Extension | Local backends | Where the answer goes | Extra Ollama setup | Source | Price |
|---|---|---|---|---|---|
| ShroomPen | Ollama, LM Studio, any OpenAI-compatible server on localhost. Also cloud keys | Into the text field you are using (insert or replace), or copy | OLLAMA_ORIGINS, once | Closed | Free |
| Page Assist | Ollama, Chrome’s built-in Gemini Nano, OpenAI-compatible servers (LM Studio, llamafile) | Sidebar and a full-page chat | None on localhost | Open (MIT) | Free |
| NativeMind | Ollama, or WebLLM inside the browser | Sidebar, with in-page writing actions | Not stated | Open (AGPL-3.0) | Free |
| Scramble | Ollama, LM Studio. Also OpenAI, Anthropic, Groq, OpenRouter | Replaces the text you selected, from the right-click menu | Not stated | Open (MIT) | Free |
| SurfMind | Ollama, LM Studio, llama.cpp, custom endpoints. Also cloud models | Sidebar | Not stated | Closed | Free for local models |
| Lumos | Ollama | Popup chat about the page | OLLAMA_ORIGINS | Open (MIT) | Free |
Checked against each project’s own site or README on 25 September 2026. “Not stated” means the project’s page didn’t say either way.
ShroomPen: a local model that writes into the page
ShroomPen is our extension, so weigh this section with that in mind. It opens next to the text field you are in, or as an overlay with Ctrl+Q, and sends your request to the provider you chose. With a local server that is your own computer. The Input Helper next to a focused field has Reply, Rewrite, Fix Grammar and Translate buttons that write the result straight into that field. In the overlay you read the draft first and then insert it, replace your selection or copy it.
You can attach context per request: the page you are on, the field, selected text, other tabs you pick (PDF tabs included), and saved Workspaces with your own notes. The same extension also takes cloud keys from OpenRouter, OpenAI, Gemini or Anthropic, so you can switch between a local model and a cloud model without changing tools.
Where it falls short: it is closed source, it runs only in Chrome and other Chromium browsers that install from the Chrome Web Store, and Ollama needs the one-time OLLAMA_ORIGINS change described below. It is not a chat app with long histories or document search.
Page Assist: the best local chat
Page Assist is an open-source sidebar and full-page web UI for local models. It works with Ollama, Chrome’s built-in Gemini Nano and OpenAI-compatible servers such as LM Studio, and runs in Chrome, Edge, Brave and Firefox. It rewrites request headers for localhost and 127.0.0.1 addresses, so on a default setup you don’t have to touch OLLAMA_ORIGINS. It says it collects no personal data and keeps everything in browser storage.
Pick it if you mainly want to chat with your model, chat with the current page or use it in Firefox. It doesn’t put answers into page fields for you. There is a longer ShroomPen vs Page Assist comparison.
NativeMind: local writing tools with a no-install trial
NativeMind is an open-source (AGPL-3.0) assistant built on Ollama. It can also run a small model (Qwen3 0.6B) inside the browser with WebLLM, so you can try it without installing anything. It offers a sidebar plus in-page actions for rewriting, proofreading and changing tone, and says no data leaves your device. LM Studio isn’t listed as a backend.
Scramble: rewrite selected text from the right-click menu
Scramble is a small open-source (MIT) Chrome extension: select text, right-click, and pick an action such as “Fix spelling and grammar” or “Improve writing”. It supports Ollama and LM Studio as well as several cloud providers, and you can add your own prompts. Its maintainer says the project has many users but is no longer actively developed.
SurfMind: a closed-source sidebar for many backends
SurfMind describes itself as a sidebar that works with Ollama, LM Studio, llama.cpp, custom OpenAI-compatible endpoints and cloud models, with export of chats to Notion and Obsidian. It is closed source. Local models are free and cloud use is pay-as-you-go, according to SurfMind’s own roundup.
Lumos: page Q&A for Ollama users who build from source
Lumos is an open-source (MIT) Chrome extension that answers questions about the page you are on, using retrieval over the page with a local Ollama model. Its README installs it by building from source or from its GitHub releases, and it needs OLLAMA_ORIGINS=chrome-extension://*.
Set up ShroomPen with a local model
- Install Ollama and pull a model, for example
ollama pull llama3.2. For LM Studio, follow the LM Studio guide instead and skip step 2. - Allow ShroomPen in Ollama by setting this environment variable, then restart Ollama:
The OLLAMA_ORIGINS guide shows where to set it on Windows, macOS and Linux.OLLAMA_ORIGINS=chrome-extension://ghigbnpoahodgnhlgbggjphpflkofgfn - Add ShroomPen from the Chrome Web Store, open it and choose “Enable ShroomPen” after reading what it does with your data.
- In ShroomPen settings choose Add provider, then Direct provider, then Local server. Enter
http://localhost:11434/v1and leave the API key empty. - Click Test/load models, allow access to localhost when Chrome asks, pick your model and click Connect.
The full walkthrough, with troubleshooting, is on the Ollama Chrome extension page.
How to choose
- You want a private ChatGPT-style chat with your local model: Page Assist.
- You want to try local AI without installing Ollama: NativeMind with WebLLM.
- You want the local model to write replies and rewrites into Gmail, LinkedIn, a CMS or a support tool: ShroomPen.
- You only need quick rewrites of selected text and are happy with an unmaintained open-source tool: Scramble.
Related
Questions
Which browser extension works with Ollama?
Page Assist, ShroomPen, NativeMind, Scramble, SurfMind and Lumos all work with a local Ollama server. Page Assist is the best known for chat. ShroomPen is the one on this list built to write into the text field you are typing in.
Why does my Chrome extension get a 403 error from Ollama?
Ollama only accepts browser requests from origins on its allow list, and chrome-extension:// origins are not on it by default. Set the OLLAMA_ORIGINS environment variable to the extension’s origin (for ShroomPen, chrome-extension://ghigbnpoahodgnhlgbggjphpflkofgfn) and restart Ollama. Page Assist avoids this on localhost by rewriting the request headers itself.
Can I use LM Studio instead of Ollama?
Yes, with any extension that accepts an OpenAI-compatible address. LM Studio serves one at http://localhost:1234/v1 once you start its server. ShroomPen, Page Assist, Scramble and SurfMind list LM Studio support.
Does a local model send my text anywhere?
Not if the server runs on your own computer and the extension sends the request to localhost. The extension itself may still have its own policies, so check what it says about telemetry. ShroomPen sends no telemetry.
Do I need a powerful computer?
Small models such as llama3.2 (3B parameters) or gemma3:4b run on many recent laptops and are enough for short replies and rewrites. Bigger models write better but need more memory and are slower.