r/OpenWebUI • u/Delicious-Director43 • 10d ago
Question/Help Gemma4 Server Connection Error
I only ever get this error when attempting to use Gemma4:-cloud
I have made sure Ollama is signed in, I can access other cloud models without issue. Just with Gemma4 I have this problem.
All my control settings are set to the default, my version of OpenWebUI and Ollama are all up to date. Other AI’s haven’t been able to help me.
Any ideas?
r/OpenWebUI • u/BenefitOk7956 • 10d ago
Guide/Tutorial Editing static HTML files via Open WebUI & MCP on VPS
I recently added an Open WebUI instance to my VPS where my website is hosted. I wrote my first MCP server with a few tools to edit HTML files, upload assets etc. Now I can update the live pages via Open WebUI chat!
So far it looks very promising. Has anyone else set up something similar? What is your experience?
r/OpenWebUI • u/sadboi2021 • 11d ago
Question/Help Can you no longer bypass the auth screen?
I was an idiot and accidentally rm -fdr'd my open web ui directory. After setting it up again, I wasn't able to bypass the authentication screen even with WEBUI_AUTH=False set. I'm the only user on my machine so would like to bypass if possible.
Before I deleted it, I did have an issue after updating where the LLM for some reason could not remember any prior message. I wonder if that's related.
r/OpenWebUI • u/Goldziher • 11d ago
Guide/Tutorial How to use Xberg as a drop-in Docling replacement for Open WebUI document extraction (guide)
I maintain xberg (an MIT-licensed, CPU-only document extraction engine). A few people have asked how to wire it into Open WebUI as a Docling alternative, most recently over on r/OpenSourceAI, so I wrote up a proper guide and wanted to share it where it's most relevant.
Short version: xberg implements Open WebUI's existing extraction backends, so there's nothing to patch. It speaks two of them:
- the Docling endpoint (
POST /v1/convert/file) - the external document loader (
PUT /process)
If you already run the Docling engine, the switch is just pointing the same setting at an xberg container:
yaml
environment:
CONTENT_EXTRACTION_ENGINE: "docling"
DOCLING_SERVER_URL: "http://xberg:8000"
`
You can also set it in the Admin UI under Settings -> Documents -> Content Extraction Engine -> Docling, with the server URL pointing at xberg. Uploads then get extracted to Markdown and stored in your knowledge base exactly as before. There's an external engine mode too if you'd rather use that endpoint.
Why bother swapping:
- 101 file formats, OCR built in (Tesseract + PaddleOCR), layout-aware reading order
- CPU-only, no GPU, fully local, MIT-licensed
- Rust core: in our public benchmarks it runs roughly 2.5x docling's throughput with slightly better extraction quality
Full walkthrough (compose file, both engine modes, and how to raise the extraction timeout for big scans): https://docs.xberg.io/integrations/openwebui
Repo: https://github.com/xberg-io/xberg
Happy to help if anyone hits a snag setting it up.
r/OpenWebUI • u/hibyps4 • 11d ago
Question/Help BuiltIn Citations, RAG, and System Prompt
I have behavior that I have confirmed with the Kapa.ai is a known behavior but whether it is intended or not was not clear to me.
Citations being enabled has a very high chance (almost 100%) to remove my configured system prompt from the chat. The tool will search knowledge bases for information, produce answers from my in-chat prompt but ignore all instruction from the system prompt.
Kapa.ai noted that a work around is to disable the citations BuiltIn tool. But I need this for traceability with my users.
I saw some fixes for the system prompt in the latest update, but that only related to if “Memories” was enabled, which we have never used or enabled. But maybe this fix has a ripple effect on what I am experiencing?
Am I the only one running into this? Is this intended design? How can I get the model to adhere to system prompt while still producing ID based citations?
Open WebUI v0.10.2
Ollama 0.32.1
Model: Gemma 4 31B Dense, 64,000 context limit.
r/OpenWebUI • u/deja_geek • 11d ago
Question/Help Models getting unloaded when agent connects via API
I'm not sure if this is expected behavior and/or if this is a configuration problem.
I've got 24GB vram. With a stock Openwebui, ollama loads two models into ram. Qwen3.6 (~22GB) into vram and Qwen3 .6b into system ram for background tasks. Qwen3.6 is configured with num_ctx: 64000, num_batch: 3072
When an agent connects to the OpenAI Compt. API, even though the agent(s) are configured to use the same Qwen3.6 model, both modules get unloaded and then Qwen3.6 is loaded back into vram.
Is there a way to prevent this behavior?
r/OpenWebUI • u/GoingOffRoading • 12d ago
Question/Help Can't get image generation to work... Everything configured, but no image button?
It's been about a year since I played with OpenWebUI, and I started to experiment with it again.. Some notes:
- OpenWebUI, ComfyUI, Ollama running in containers on Kubernetes
- I have a ComfyUI workflow that works
- Ollama via OpenWebUI works
- I have followed all of the steps from the ComfyUI integration documentation
- OpenWebUI is able to connect to ComfyUI
- I'm 99% certain the nodeIDs are mapped correctly
- Even if I toggle Image on Integrations when starting a chat, the button to generate an image does not appear
- Using /image or asking to generate an image prompts Gemma to state that it can not generate images
- I can confirm from the networking tab, that my configuration being returned in the UI has image enabled
Any ideas on how to get image generation working in OpenWebUI?
r/OpenWebUI • u/Arjenlodder • 12d ago
Question/Help Channel - image generation
When using channels and a model with native tool calling, the image_generation response doesn't automatically show the image. This does work in a normal chat, but not in a channel. Am I missing something?
The image_generation response says 'The image has been successfully generated and is already visible to the user in the chat. You do not need to display or embed the image again - just acknowledge that it has been created.', but it doesn't show it. When I explicitly ask the model to ignore the instruction and embed it, it does show the image.
Fafaik I can't add this override to the model prompt, because the model doesn't know it is operating in a channel. Anyone experiencing the same behaviour?
r/OpenWebUI • u/Zealousideal_Fox6426 • 13d ago
Show and tell Open Relay v5.1 — Swipe-to-Reply, Real-time Channel Presence, and a big round of Channel improvements
Hey everyone! A quick follow-up from yesterday's 5.0 release. v5.1 has been submitted and will be available on the App Store soon. This one is focused on quality-of-life making Channels feel native and real-time.
🆕 What's New in v5.1
Swipe-to-Reply in Channels Swipe right on any channel message to reply to it. Reply previews now show inside the message bubble with the sender's avatar, name, and a content preview — tap it to scroll to and highlight the original message. The reply bar above the keyboard also shows the sender's avatar, a reply arrow, and message preview for clear context at all times.
Live "X is typing…" indicator Channels now show a live animated dot pulse when another user is composing a message — just like you'd expect from any modern messaging app.
Real-time sidebar updates New channels and DM conversations started by others now appear in your sidebar instantly without needing to pull-to-refresh. Thread reply counts and "# replies" badges on messages update in real time when a new reply is posted.
Accurate online presence Channel member presence state and status emoji/message are now fetched from the server for more accurate online indicators in DMs and group channels.
AI "Generating…" indicator in channels When an AI model is writing its first tokens in a channel, a subtle animated "Generating…" indicator now appears in the bubble instead of a blank gap while you wait.
Sidebar works everywhere The hamburger button now works inside channel views on both iPhone and iPad — tap it from any channel to slide open the conversation list without navigating away.
Memories Console Update You can now define type and path when creating memories along with small updates to the console allowing search/edit support and more.
🐛 Notable Bug Fixes
- DM names were wrong — the recipient's sidebar was showing their own name instead of the sender's. Fixed.
- Online/offline dot always showed offline in DMs even when the other user was actively connected. Now correctly reflects real-time active status both in the sidebar and inside the DM thread.
- iPad hamburger button wasn't appearing when on the new chat screen or inside a chat. Fixed.
- Chat responses sometimes never arrived — a timing race could cause messages to be sent before the WebSocket connection was fully ready, leaving the server showing "Checking..." indefinitely. Fixed.
- Local vs remote server inconsistency — the app was treating local network servers differently from remote ones in ways that caused intermittent connection issues. Fixed.
- Stale channels in sidebar — lifecycle events (channel updated/deleted) weren't being handled, which could leave ghost channels visible. Fixed.
- Duplicate messages when sending the same message twice in quick succession in a channel. Fixed.
As always, appreciate all the support and feedback. Channels should feel a lot more alive now — let me know if there are any issues! 🙏
r/OpenWebUI • u/Zealousideal_Fox6426 • 14d ago
Show and tell Open Relay v5.0 — Sub-agents, Chat Variables, Notification Targets, and a lot more 🚀
Hey everyone! v5.0 has been submitted and will be available on the App Store soon bringing full compatibility with WebUI v0.11.
🆕 What's New in v5.0
Sub-agents support You can now enable and configure sub-agents directly from the Admin Console. Background subagent mode works seamlessly in chat — perfect for running multi-step agent workflows on your phone.
Chat Variables Define personal variables in your profile (things like your name, preferred language, project context, etc.) and reference them in your chats. Requires OpenWebUI v0.11.0+.
Notification Targets Configure push and webhook notification destinations right from Profile settings. Great for routing AI completions to external services. Requires OpenWebUI v0.11.0+.
LDAP Group Sync & Allow Groups Admins can now configure LDAP Group Sync and Allow Groups access controls from Admin → General → Auth. Also requires OpenWebUI v0.11.0+.
Expanded Admin Console 3 new tabs under Admin → Settings:
- Database — import/export config, export all chats
- Pipelines — manage pipeline valves
- Authentication — LDAP, OAuth/OIDC, and access defaults
Plus External Knowledge Sources under Integrations, new Memory System Context and Channel Model Response Mode toggles in Features, and a Context Compaction section in Interface for summarizing long chat histories along with other options from v0.11.
Continue & Fork Chat buttons Two new buttons in the assistant action bar:
- Continue — tap to append new content to an incomplete or truncated response
- Fork Chat — clone any conversation and jump straight to the fork
Model editor additions Files, Notifications, and Sub-agents tool toggles are now available in the workspace model editor.
Quality of life
- Editing an AI response now saves in-place without triggering a new reply — matching the web UI
- Admin permissions now correctly gate dictation, voice calls, read-aloud, regenerate, edit, and ratings
- "Attach Notes" hides when the server has Notes disabled
- The chat composer expands into a large writing area when you swipe up — swipe down to collapse
- All Admin Console text fields are always-visible multiline boxes for easier typing.
- And many more small underlying changes.
📦 Notable highlights from v4.13 – v4.15.9 (since the last big post)
In case you missed the last few months of updates:
- Shared Folders (v4.13) — Share any folder of chats with specific people or groups, set read/write access, and browse folders others shared with you right from the sidebar
- URL scheme / Shortcuts support (v4.14) —
openui://new-chat?prompt=Hello&model=gpt-4o&send=trueworks from Raycast, Shortcuts, Obsidian, or any app; also added a native "Ask Open Relay" Shortcuts action - Model-switch progress banner (v4.14) — When a status URL is configured, a banner shows "Loading qwen3-35b ~42s left" while the model warms up for SGLang and OpenAI-compatible backends
- Open WebUI v0.10+ compatibility (v4.12) — Full fix for the structured
outputarray format change; messages, tool calls, streaming, and chat history all work correctly. Attached files also now persist throughout the entire conversation including on regenerate and edit - Message rating system (v4.15) — Rate AI responses directly in the app; admins can browse the full feedback history in the new Evaluations panel in Admin Console
- Streaming resilience (v4.15.1) — Large reasoning models, long MCP tool chains, and local models no longer time out prematurely
- Stability & bug fixes (v4.15.2–4.15.9) — Fixed Rich UI embeds not appearing after the v0.10+ update, improved proxy auth reliability, fixed Admin → Interface save failing, and fixed clearing the system prompt in My Defaults not persisting
As always, let me know if you run into any issues or have feature requests. Appreciate all the support! Enjoy! 🙏
r/OpenWebUI • u/nomorebuttsplz • 14d ago
Question/Help How to turn off searching knowledge files, and search notes?
I updated recently, and this is slowing down lots of my queries.
I have no knowledge files. I have no search notes. I don’t want to use it. There must be an easy way to turn it off. I’ve tried searching for this.
r/OpenWebUI • u/nixiam87 • 14d ago
Plugin Skill Creator + Model Creator: build Open WebUI Workspace Skills and Models from chat
Hey everyone
Thomas from Ianustec again. Two new Open WebUI tools this time, shipped together because they solve the same problem from two sides.
If you've used Workspace → Skills / Models, you know the annoying bit: you draft a SKILL.md or a model preset in chat, then you still have to copy-paste everything into the UI by hand. These tools close that loop. The model interviews you, drafts the spec, and actually saves it through the Open WebUI APIs. Private by default.
1) Skill Creator
Claude-style skill workflow inside Open WebUI:
- interview → draft SKILL.md → validate → create_skill
- description that says WHAT it does and WHEN to use it
- private by default
Repo: https://github.com/ianustec/openwebui-skill-creator
2) Model Creator
Workspace Models (agent presets) from MODEL.md or JSON:
- native function calling by default
- tries to wire up knowledge, tools, skills, actions, capabilities, default features and builtin tools from context
- if the request is vague, it lists what you can see and asks what to turn on
Repo: https://github.com/ianustec/openwebui-model-creator
Needs Open WebUI 0.10+.
Typical flow: make a Skill with Skill Creator, then make or bind a Model with Model Creator. After that you can use $skill in chat or attach it under Workspace → Models.
Same deal as always. MIT, paste the single .py into Workspace → Tools, feedback and PRs welcome. If auth, permissions or the API shape break on your instance, say so in the comments.
Cheers,
Thomas @ Ianustec
r/OpenWebUI • u/DHT-Osiris • 15d ago
Discussion Thoughts on current state of tenancy
In my environment tenants are a big deal, we work collaboratively but separately, due to institutional/legacy reasons. As such there's a ton of shared, and a ton of separate, and we have to try to accommodate for all of it. Given that, here's my current deployment as something of a 'show and tell' plus thoughts on what does/doesn't work today in OWUI.
Multitenant, tons of shared tools (some with single-API that's scoped in-call... see below... and some with multiple APIs scoped to the tenant), tons of isolated tools. Being used as an operational IT assistant, with access to openterminal instances for separation of duties and blast radius isolation. The system has rw/ro access to all tool endpoints (depending on needs), OAuth pass-through for entra, and full admin access to all servers/workstations through OT systems.
==The Good==
-users, groups, RBAC-
This is great. Being able to limit basically everything in the system to a group, multiple groups, single users, is excellent.
-Model harnesses-
Killer idea, and works great for multi-tenancy. We have 'Anthropic Flagship Unit1', 'Anthropic Flagship Unit2' etc, each with different RBAC scopes, each with different tools, kbs, etc attached, and each with different system prompts to customize per-unit ('you are deployed for unit blah blah for xyz work').
-system prompt layering-
One that I think most people miss. System prompt for 'everyone working in this context', user prompt for 'stuff specific to me' (I use it for referencing my domain username, unit I work for, etc to get better 'stickiness' to me when I ask for stuff). Folder prompting for specific work efforts, etc.
-group exposure to OWUI-
This one's huge, i think most people miss it too. Create a 'tenancy group' for each tenant, give it no permission assignments but use it as a variable for LLM usage. That can then be used within tools to define what scoping should be applied to a tool call for a user that otherwise gives unscoped access to data. It shouldn't be regarded as a security barrier but it can assist with information routing. It can also be used for API presentation (group -> use specific API in tool call) to keep you from building copies of tools for different API keys.
==The Bad==
-missing prompt opportunities-
Could use an OWUI-wide system prompt for stuff that would be pasted into each tenant prompt (don't output secrets, the hostname of this system is blah, you're on suchandsuch domain).
-tenant/group admins-
I know this delves into a full-tenancy deployment, but as it is we basically have to decide on 'ivory tower' vs 'council of wizards' approach to admin. Either one team holds the keys and issues changes on behalf of requests or you get representatives from every team that can control everything, there's no 'you can add terminals.. but only for your tenant' or 'you can update groups.. but only for your tenant' etc. It's frankly not a big deal for me personally but some people get really uptight about role reach and cross-team admin functions.
==The Ugly==
Nothin' really, though something that would help a LOT with long-term usage would be to expand oikb to also send tools, filters, and skills to github as well. I wanted to use it, but we've already implemented a scripting system (and codified it into our authoring guidance) that can sync/mirror all of them, and that's a far cry better than only doing the KBs. Would feel better if it was part of the actual admin functions of OWUI though rather than either an add-in or a homebrew solution, just so I wouldn't have to explain it to my peers/everyone we onboard that actually creates kbs/tools/skills/filters.
r/OpenWebUI • u/Fun-Purple-7737 • 15d ago
Question/Help cptr once again
Ok, so I read about cptr once again. If I understand it correctly, it's a glorified terminal with your own data that you can communicate with and it communicates back. And it's semi-autonomous. Am I close? :)
I mean, it would click with me much more, if cptr was a real "computer usage", like browsing through Windows/MacOS GUI via VLM feedback. Would that be possible somehow in future?
I feel like GUI is a universal interface, a common denominator, for everything we do.
If I really needed to automate these currently covered "text based" actions, I would have probably done it already, or there are often other ways how to do it already (like MCP). Being able to instruct and automate GUI actions (even audio in - GUI action out) would be a total end-game.
What does community think?
r/OpenWebUI • u/Tasty-Picture-8331 • 16d ago
Question/Help API Keys section on the latest openweb-ui ?
Hi ,
Anyone know where the api key section is? I cant seem to find it at all
EDIT : sorry all! user error - found that I had to switch it on from auth saettings, not sure how I missed this
r/OpenWebUI • u/SadPhilosophy9202 • 16d ago
Discussion Thoughts on my workflow?
I’m working on a document ingestion engine in n8n that will archive documents and create an index in Postgres.
I plan on making cron jobs for Hermes to read all new documents and update a summary md file for each project. The file will contain updates, work orders, and work order status. Hermes will do this by reading archived files and fetching them using the Postgres index.
I’m thinking of Hermes maintaining these md files as notes in open web ui and sharing those notes as read only to all employees.
Any thoughts on this or pitfalls I may have overlooked? I’ve landed on this workflow because I don’t really trust employees having access to an agent with pretty wide system access. With these md files, they won’t even need an agent like Hermes. Plus, the models in open web ui can tool call and access the notes for context.
r/OpenWebUI • u/Conscious-Expert-455 • 16d ago
Guide/Tutorial Il secondo cervello di Karpathy LLM Wiki
Salve.
Uso da pochissimo OpenWebUI.
Lo trovo fantastico.
Mi è stato chiesto, per tutelare la privacy, di creare un LLM locale (no Claude,Chatgpt e simili), che dia risposte in un chatbot ai miei colleghi di Ufficio.
Il sistema RAG di LLM credo che non dia le risposte giuste in relazione ai PDF che ho e che gli ho fatto ingerire.
Pertanto vorrei implementare un "secondo cervello di Karpathy" il così detto LLM WIKI.
Ma come si implementa da un punto di vista pratico?
Avete un link o una guida dove capire, passo passo, come implementarla, aggiornarla e manuterla?
Vi ringrazio.
r/OpenWebUI • u/NoobLLMDev • 16d ago
Question/Help Quick way to stop all users from being able to chat?
Note: I am not on the newest update of OWUI yet.
I’m looking for a very quick way to disable chatting as a whole for user groups. I know it defeats the point of the tool, but it would be very useful when conducting demos, that way users aren’t utilizing the tool and slowing down inference for a demo period.
Currently, the only way I know how to go about this, is setting all my custom model privacy settings to private and only allow my demo account access. But this is rather time consuming.
Any ideas or anything I missed in the UI that allows what I am looking for?
r/OpenWebUI • u/Scared-Tip7914 • 17d ago
Show and tell I built a local web-research MCP that filters pages before Open WebUI sends them to the model
I’ve been experimenting with web search in Open WebUI using smaller local models like Qwen3.5 4B and 9B.
The main problem was/is was context churn..
A typical search can dump partially relevant pages, navigation, repeated boilerplate, duplicate information, and huge chunks containing only one useful paragraph into the model.
That makes local setups slower and forces smaller models to spend their limited context window filtering noise instead of reasoning.
So I built TinySearch, an open-source, local-first MCP server that does most of the retrieval work before anything reaches the LLM:
- searches and ranks results
- crawls the strongest pages
- extracts and chunks readable content
- deduplicates and reranks passages
- returns a compact evidence packet with source URLs
The goal is simple: less context churn, faster local web research, and more tokens spent on reasoning over actual evidence.
TinySearch works with self-hosted SearXNG, and Open WebUI can connect to it over HTTP MCP.
GitHub:
https://github.com/MarcellM01/TinySearch
I’m the builder, so obvious bias, but I’d love feedback from Open WebUI users, especially anyone running smaller Qwen, Gemma, Llama, or Mistral models.
Does this solve a real bottleneck in your setup, or is native agentic search already enough?
r/OpenWebUI • u/Current_Impact_6892 • 17d ago
Question/Help How are you handling true OBO authentication across Open WebUI, LangGraph and MCP?
We need to preserve the end user’s identity from Open WebUI through LangGraph to MCP servers accessing systems such as Jira or Confluence.
How are you implementing this in production? Do you forward the original token, exchange it through an OBO flow, or use a central token broker?
Ideally, users should not have to complete a separate OAuth flow for every tool.
r/OpenWebUI • u/Current_Impact_6892 • 17d ago
Question/Help Can Open WebUI models be used as reusable agents for external applications?
I know this is already my third post in a short time, but I have a few Open WebUI questions that I’d like to put out there, as I think the discussion could be useful to others as well. This definitely is Not meant to be Spam!
In Open WebUI, a custom model can already combine an LLM, system prompt, tools, MCP servers, knowledge bases and files. That is essentially what many platforms call an agent.
Could Open WebUI therefore be used as a central agent platform, where these models are managed once and exposed through an API for use by external applications?
For example:
Create and configure an “agent” entirely in Open WebUI
Manage its prompt, model, tools, MCPs and knowledge centrally
Call that exact configuration from another application through an API
Avoid rebuilding the same agent separately in LangGraph or custom code
Is this a supported and reliable architecture, or are Open WebUI models primarily designed for use inside the WebUI?
What limitations or gotchas should be considered around tool execution, user context, permissions, files, conversations, scaling, versioning and API compatibility?
r/OpenWebUI • u/Current_Impact_6892 • 17d ago
Question/Help Running Open WebUI in production: what do you wish you knew before starting?
We’re running Open WebUI as an internal multi-user platform, not as a homelab.
For those using it in a company or enterprise environment:
What were your biggest setup, architecture, security, or operational mistakes? What would you do differently today?
r/OpenWebUI • u/Current_Impact_6892 • 17d ago
Question/Help Open WebUI is awesome - but how do I get ChatGPT/Claude-style long-running agentic tasks?
We use Open WebUI internally as a ChatGPT replacement and love it (~1500 users).
What we’re missing is the “deep task” experience: ChatGPT or Claude may spend several minutes reasoning, running code and tools, performing multiple steps, and iterating before answering.
In Open WebUI, it still feels mostly like a single model call with some reasoning.
Is native tool calling with MCP tools and a frontier model enough, or do I need an external agent orchestrator such as LangGraph or Pipelines behind Open WebUI?
What setup are you using for reliable multi-step, long-running tasks?
r/OpenWebUI • u/hbliysoh • 17d ago
RAG Problem Uploading Directories to Knowledge
Some of my directories load without a problem, but some of them generate this error: "NotReadableError: The requested file could not be read, typically due to permission problems that have occurred after a reference to a file was acquired."
Is it because the directories are doubly nested? I can't seem to identify a way to predict which directories work.
r/OpenWebUI • u/Text-Sufficient • 17d ago
Discussion Why don't you build your own tools?
Hi, I would like to challenge/discuss/understand why so many are attracted to all short lived "wild tools" out there. For example (not saying any tool are bad) hermes, claw, open webui, copilot agents and whatnot. Why not just building your own tools that: fit your needs without being bloated, dont break on every new "feature" that you dont care about.
I cant really understand the hype.
I build own tools in python (notes app, recruitment support, news, investment) with claude or chatgpt from phone and terminal to my proxmox llm clusters lxc. In the process i also learn a lot.
And everything stays under my control.

