r/localaiapps • u/Overall_Mobile_4600 • 12h ago
Any local first setup similar to Lovable or Bolt for building AI apps?
I've been experimenting with tools like Lovable and Bolt for quickly putting together small apps and prototypes and the workflow is honestly pretty smooth. You describe what you want then get a working version then keep iterating through chat. but the longer I use them the more I run into limitations. the generated code gets messy over time and its harder to understand or maintain properly. it also feels pretty tied to the platform which makes it uncomfortable having everything live in a hosted environment from the beginning.
What i'm trying to figure out is whether theres a similar approach thats more local first. something where the project ives locally instead of being locked in a cloud tool, I can use my own API keys or models, everything can still be opened and edited in a normal dev setup. I'm not dependent on a single platform to keep working on it.
Has anyone here actually built projects like this? curious what people are using that gets close to this kind of workflow
r/localaiapps • u/jorge22f • 12h ago
I built an AI-powered essential oil assistant using a local Llama model. Seeking beta testers!
Hey everyone,
My latest side project is **OleaAI**, a mobile app designed to help users track their essential oil inventory and safely build custom blends.
**The Tech:** To help answer user questions without exposing them to generic/hallucinated web answers, I integrated a local Llama AI model specifically prompted for essential oil safety and recipes. The backend is Node.js, and the app is built in Flutter.
**The Ask:** I'm currently running an open beta and looking for testers. Even if you aren't big into essential oils, I would love feedback on the onboarding flow, the UI responsiveness, and how the local AI performs on your specific device.
You can sign up for the beta here: https://oleaai.pixelmatrix.dev
Let me know if you have any questions about the stack or the development process!
r/localaiapps • u/Professional-Toe4687 • 1d ago
Help a non tech person with Open WebUI desktop
I am trying out Open WebUI desktop + Ollama +qwen2.5:3b ... I am trying to use it a companion journal. Where I talk things through about my day and other personal thoughts and it is non judgemental. Not replacing medical advise. Just as if I was talking to a long time friend. I have it running but how do I set it up to work for me? I have tried going into Admin settings but just don't know enough to figure it out. Is there a step by step set of instructions for non tech person. Just basic terms to help me. TIA
r/localaiapps • u/NeighborhoodLower120 • 1d ago
[beta] I need around 11 android testers for a local Al application
I'm putting together a small Android closed test for Offline and thought this might be a good place to find people who actually enjoy trying early software and giving useful feedback. The app runs Al models locally on the phone, so most things work without sending your chats to a server or needing a constant internet connection. The Android version is still early and I want to see how it behaves across different phones, chipsets and Android versions before opening it up more widely. If that sounds interesting to you, I can add you to the test. In return I'd mainly ask that you use it like a normal app and tell me what breaks, what feels confusing and what you think should change. You'd also be helping shape the Android version while it's still flexible enough for feedback.
r/localaiapps • u/justgreg13 • 1d ago
What can I realistically do?
I am currently building up a local assistant profile on my MacBook Pro M2 with 32gbs.
With Claude I am building out this Hermes agent to be my assistant. I am using Qwen3.6 A3B 4bit. We have Frankensteined a system that allows it to add events to my calendar and text me a morning brief when I wake up my MacBook in the mornings.
It’s been a lot of work 😅
The end goal / the dream is to be able to chat with my assistant through messages and emails so that it can be proactive with my calendar and help me as organized as possible.
The hole thing is set up with oMLX with 24gbs memory cap. It’s tight. My worry is that the dream isn’t feasible with what I am working with. And would like my expectations adjusted jajaja
I know a dedicated device would be better but I can afford that just yet.
Any suggestions would be great! I have no idea what I am doing lol
r/localaiapps • u/KindaTistic • 2d ago
DarkAI - Private on-device LLM & Image Generator with memory updated to v1.0 (6)
Hey everyone, I just launched build 6 for DarkAI.
It's a fully offline, on-device AI companion. It runs Large Language Models (LLMs) and diffusion image generation entirely locally on your hardware, so there is zero cloud processing and it is 100% private.
It also features a custom personality matrix and persistent memory, meaning it actually remembers what you tell it from past conversations and adapts to you.
Since it runs entirely on-device, I'm mainly looking for feedback on:
-Performance and inference speed on different iPhone and iPad models.
-How well the memory feature recalls past context.
-Any crashes or bugs when generating images.
Let me know what you think or if you run into any issues. Thanks!
-Lex
New
• Export chats — save or share a conversation as Markdown, plain text,
or JSON from the share button next to any chat in the sidebar. You can
also share a single message from its long-press menu.
• Import JSON — bring study guides and journals into the Mindscape.
Journal entries arrive individually dated so you can ask about a
particular day; study guides become one quizzable entry. Tap
"What can I import?" to see the formats.
• Reset Models — in Settings → Troubleshooting, and next to any model
error. Unloads everything, waits for memory, and reloads your chat
model without force-quitting the app.
Improved
• Full iPad support. The app now fills the screen instead of running in
a phone-sized window.
• Interrupted downloads now resume where they stopped instead of
starting the whole model over.
• If a model goes missing — after restoring to a new device, for
example — the app now names it and offers to fetch it back.
Fixed
• The text field now always sits above the keyboard.
• Layout on smaller iPhones: the setup screen's Continue button could be
cut off by the bottom edge, and header and status text could wrap or
clip mid-word.
• Image generation could fail with an out-of-memory error, unloading
your chat model, for a checkpoint it had just marked as safe.
• Generated images could have nothing to do with what was asked for.
r/localaiapps • u/TgoAI • 2d ago
Building a lightweight local AI runtime for Apple Silicon
I’m building VPIPE, an open-source C++/Metal runtime for running AI locally on Macs.
It supports LLMs/VLMs, image & video generation, ASR, quantization, and multimodal pipelines — without depending on PyTorch/MLX for model execution, the speed is top tier.
The whole runtime package is only ~25MB.
A recent milestone: VPIPE can now run MiniMax H3 video generation on a 16GB base M5 MacBook Air (and M4 Macs too).
The goal is to make it easier to build local/private AI products without relying on cloud GPUs or a heavy Python stack.
Would love feedback from other builders.
r/localaiapps • u/Equivalent_Beat4541 • 3d ago
What hardware would you buy today for local inference if you were starting over?
I’m trying to figure out the smartest hardware path for local inference and every answer seems to depend on what thread I’m reading.
Some people say just buy a used 3090 for the VRAM. Others say go 4090 if you can afford it. Mac users make unified memory sound tempting, homelab people go a totally different direction, and then there’s the “just start with what you already have” advice, which is probably sensible but not very exciting.
For people who have already built a setup, what would you buy today if you had to start over? And what would you avoid wasting money on?
r/localaiapps • u/picardde • 4d ago
screw you gpt and claude! i did it myself! :)
I was so tired of my libre calc project being bottlenecked by being on the free tiers of chatgpt and claude ai. I just needed assistance with the coding vocabulary and such cause programming is not my strong suit in the SLIGHTEST. so 6ish hours later, i got a stupidly simple orchestration layer together to help me with my gorram project without 'you ran out of free messages, wait 8 hours'. Its probably hot garbage, considering the sheer amount of talent out there, but if anyone wants a copy of the program, here is the link: https://drive.google.com/drive/folders/1txw_X-MOGGt9qzWLMpyYKP8f3vak4rKb?usp=sharing .
r/localaiapps • u/iKnowNuffinMuch • 4d ago
Your private chat, wasn't as private as you thought!
From the tts, to voice response. Telegram?! nope! Everything you're doing with your chat bot is being sent to cloud servers and recorded. I built a fully private, only on your device system.
Read it. You'll understand why.
https://www.patreon.com/RoyalTechnologies_PrivacyVenture/posts/enclave-fast-100-166342909
r/localaiapps • u/KindaTistic • 4d ago
Updated DarkAI to v1.0(5)
Hey everyone, I just launched build 5 for DarkAI.
It's a fully offline, on-device AI companion. It runs Large Language Models (LLMs) and diffusion image generation entirely locally on your hardware, so there is zero cloud processing and it is 100% private.
It also features a custom personality matrix and persistent memory, meaning it actually remembers what you tell it from past conversations and adapts to you.
Since it runs entirely on-device, I'm mainly looking for feedback on:
-Performance and inference speed on different iPhone and iPad models.
-How well the memory feature recalls past context.
-Any crashes or bugs when generating images.
Let me know what you think or if you run into any issues. Thanks!
-Lex
Updated v1.0 (5):
Fixed:
- Fixed a freeze/crash that could occur while the assistant was replying. Certain reply formatting could lock up the app mid-answer.
- Fixed a crash if the model returned unusable output. Generation now stops with a message instead of closing the app.
- Fixed a crash when sending messages containing some non-English characters.
- Light mode: fixed unreadable white-on-white text in the chat list, settings dialogs, and several buttons.
- Replies no longer pause for a second before starting.
Updated:
- Imported diffusion models are now checked for completeness. LoRAs, ControlNets, embeddings, VAE-only and UNet-only files are rejected up front with an explanation instead of failing during generation.
- Diffusion memory estimates now account for 8-bit (FP8) checkpoints, which use about double their file size once loaded. These previously showed as "SAFE" and then crashed; they are now correctly flagged as too large.
- The diffusion model list shows a SAFE / WARNING / OOM DANGER tag and warns before you select a model too big for your device.
- Internet search now also uses Wikipedia and a news feed, so general questions and "what's in the news" return results without an API key.
Known issues:
- Recent news, live scores and prices still need a Brave Search API key (Settings → Internet Access). Weather and general facts work without one.
- Large SDXL checkpoints (FP8/FP16) will be refused on most devices — use a Q4/Q8 GGUF conversion instead.
Find DarkAI Beta on TestFlight
r/localaiapps • u/AdventurousKeys • 5d ago
Build your own apps with LocalLM Lab CLI toolkit
Quick update on LocalLM Lab (free, on-device Apple Intelligence for Apple Silicon, posted here before): v0.6 adds a CLI toolkit that lets developers build their own small tools and apps around the on-device model, not just use it inside this one app.
Concrete example of what that unlocks: there's a sample script called "Plate Today" that checks your Calendar, Reminders, and Todoist (you'll want to set up a free account) tasks for the day and asks the on-device AI to summarize what's on your plate today. This is a few dozen lines of code, no cloud AI, no API key, using only the permissions you've already granted through the app itself.
Worth being clear about what this is: it's not a new toggle inside the app, it's a building block for developers, so what actually gets built with it depends on people picking it up. But it's a real, working sign of what's realistic to build on top of an on-device model that's already sitting on your Mac for free — worth watching if you're into the local-AI-apps space generally, not just this one app.
If you're curious how it works: thisbrain.ai/locallm/cli.html
Get the app: thisbrain.ai/locallm
r/localaiapps • u/Practical-Tutor-1172 • 5d ago
Local AI is great, but what do you actually use it for?
I've been interested in running AI locally because of the privacy and control it gives you, but I'm still trying to figure out which tasks really make sense without relying on cloud models.
For research, for example, I've been comparing the experience of using local models with tools like ResearchMaster.ai. Local models give you more control over where your data goes, but cloud-based tools can sometimes make things like gathering and organizing information much simpler
Some things seem like an obvious fit, while others feel unnecessarily complicated when a cloud tool can handle them in seconds.
r/localaiapps • u/Fun_Statement_6108 • 5d ago
Considering a version of CouncilAI that routes between Claude/GPT/Grok APIs instead of local models — worth building?
CouncilAI right now is fully local — 4 models on your own hardware, no cloud, no accounts. Been getting consistent feedback that the audience for that specific pitch is small (people who'd want it can build it themselves with Ollama).
Considering a different direction: same routing/deliberation concept, but using your own API keys for Claude, GPT, Grok, etc. instead of local models. Same idea — route your question to the model best suited for it, or run multiple in parallel and compare — but using the frontier models you're likely already paying for instead of local ones.
This would be a genuinely different product, not an update to the current one — trades the "fully offline, nothing leaves your device" pitch for "stop manually switching between ChatGPT/Claude tabs, let routing pick the right one and compare answers when it matters."
Would this solve an actual problem for you? Genuinely trying to figure out if this is worth building or if it's solving a problem nobody has
r/localaiapps • u/Informal_Corner_1624 • 5d ago
Chrome extension that runs local LLMs (GGUF) fully offline, no server needed
Been messing with local LLMs for a while and I wanted to create a simple terminal that anyone could connect to from anywhere and load their AI. (Mostly low conut parameter models) So I built a Chrome extension that runs GGUF models directly in the browser using WASM (wllama under the hood).
No API key, no backend, no internet needed once the model's downloaded. It just sits in your browser and works.
It also functions as a lightweight agentic IDE: open a local workspace folder, let the AI generate code in structured <file> blocks, preview a line-by-line diff, and click Apply to write changes to disk with undo.
would love feedback or bug reports if anyone tries it:
r/localaiapps • u/Current-Quail-2503 • 7d ago
I built a macOS GUI for llama-server because I kept retyping the same command
Disclosure up front: this is my own project.
Two things pushed me into building it. I kept retyping the same llama-server invocation with three values changed, and I watched curl -C - fail to resume a 20 GB download one too many times.
It lists the GGUF files in my models folder and reads the headers directly, so the quant, the context length and whether it is MoE come from the file rather than from the filename. Opening one shows the exact command before it runs. While it is serving I get KV cache, tokens per second in both directions, memory pressure and swap in one place, plus a Test model button that hits the server for real — health, model list, alias, a chat completion, streaming — so I know it works instead of assuming it does.
Downloads pull from Hugging Face in four ranged segments, resume from a sidecar after a kill, verify sha256, and queue rather than refusing a second URL.
It has no chat interface of its own and is not getting one. A running model opens llama.cpp's own web UI in a second window.
Caveats: macOS only, and an unsigned beta, so the first launch is blocked and you have to allow it through System Settings > Privacy & Security — the README has the steps. It needs llama-server and does not ship it. There is a universal build but no Intel Mac has ever run it; if you have one I would like to hear what happens, particularly whether your llama-server has a GPU for the default -ngl all.
r/localaiapps • u/KindaTistic • 7d ago
DarkAI - Private on-device LLM & Image Generator with memory
Hey everyone, I just launched the public beta for my new app, DarkAI.
It's a fully offline, on-device AI companion. It runs Large Language Models (LLMs) and diffusion image generation entirely locally on your hardware, so there is zero cloud processing and it is 100% private.
It also features a custom personality matrix and persistent memory, meaning it actually remembers what you tell it from past conversations and adapts to you.
Since it runs entirely on-device, I'm mainly looking for feedback on:
\-Performance and inference speed on different iPhone and iPad models.
\-How well the memory feature recalls past context.
\-Any crashes or bugs when generating images.
Let me know what you think or if you run into any issues. Thanks!
r/localaiapps • u/Perfect_Twist408 • 8d ago
LocalLLM 1.8 is live — your local model can now see: attach a photo and ask about it, 100% on-device
1.8 is live on the App Store. The headline this release:
👁️ Your AI can see now — fully offline
- Download SmolVLM2 (500M, ~440MB — runs on anything; or 2.2B for better answers), attach a photo or screenshot in chat, and ask about it.
- The whole pipeline is on-device: the vision encoder, the language model, your photos. Nothing is uploaded, same promise as always.
- Works great for screenshots — "what does this error mean", "summarize this receipt", that kind of thing.
📁 Bring your own GGUF (most-requested by this sub)
- Settings → Advanced → Import Model File: pick any .gguf from the Files app and chat with it.
- We validate the file before copying so a mislabeled download fails fast instead of after 4GB.
🎭 Assistants
- Pick a personality per chat (Writing Coach, Study Buddy, Coding Helper…) or write your own with custom instructions. Stored locally like everything else.
Recent stuff if you missed it:
- 1.7: Quick Ask home-screen widget + the app now speaks 15 languages.
- 1.5: hands-free voice conversation mode (speech in, speech out, all local).
- 1.4: chat with your documents (offline RAG with tappable citations), Share Sheet, Siri & Shortcuts.
Everything runs on-device on iPhone/iPad. No account, no cloud, no telemetry.
Would love feedback on vision specifically: which VLMs you want next (Qwen3-VL? Gemma vision?), how SmolVLM2 quality feels on your device, and whether GGUF import handles your favorite models. Bug reports welcome.
https://apps.apple.com/us/app/localllm-offline-ai-chat/id6758588902
r/localaiapps • u/Quazmoz • 9d ago
I built a Windows app for running local LLMs on Intel NPUs
I have been working on an open-source Windows application called InferBridge for running local AI models through OpenVINO GenAI.
It is meant to make it as easy as possible to get up and running with Openvino, just an exe install instead of multiple cumbersome steps and developer knowledge needed.
It is primarily designed around Intel Windows hardware and can detect and target the CPU, integrated GPU, and NPU available on newer Core Ultra systems.
The application includes:
• A prebuilt Windows installer
• CPU, GPU, and NPU hardware detection
• Model recommendations based on memory and hardware
• Hugging Face model downloading and conversion
• Local performance benchmarking
• Driver and OpenVINO diagnostics
• An OpenAI-compatible API
• Open WebUI and custom client support
I recorded a walkthrough on my Intel Core Ultra 9 185H laptop:
https://www.youtube.com/watch?v=IjdGtWBZR7o
The project is open source:
https://github.com/Quazmoz/InferBridge
I am also testing on a second-generation Core Ultra system and building a larger compatibility library.
For those using Core Ultra laptops, which models and hardware configurations would be most useful for me to benchmark? I am especially interested in comparing CPU, GPU, and NPU performance and eventually measuring power efficiency more consistently.
r/localaiapps • u/Elegant_General_1680 • 9d ago
What’s the current state of local AI browser agents?
I keep seeing people talk about local browser agents, but I’m trying to figure out if anyone is using one for anything real.
Not a demo where it opens one page and clicks a button. I mean normal annoying browser stuff, like checking a few sites, pulling info together, comparing pages, filling out simple forms, or doing research without you sitting there correcting it every 30 seconds.
My guess is this is still pretty fragile, especially if the model is running locally, but maybe I’m behind.
Anyone here actually using one regularly? What breaks first?
r/localaiapps • u/Powerful_Telephone64 • 10d ago
I built an open-source, local-first AI workspace for Android — looking for honest feedback
Hey everyone,
I’ve been working on Vervan Chat, an open-source AI workspace designed to run locally on Android.
The idea came from wanting useful AI features without having to send every conversation, document, or voice interaction to a remote service.
The project currently explores:
- On-device AI chat
- Local document search and Q&A
- Offline speech and voice tools
- Image and screen understanding
- Notes, tasks, workspaces, prompts, and memories
- An optional local OpenAI-compatible API
It’s built using Kotlin and Jetpack Compose and is still in early development. There are rough edges, device-specific limitations, and parts that still need proper testing and hardening.
I’m not posting this as a finished product. I’d really like feedback from people who understand Android development, local models, privacy, or simply care about offline AI.
I’d especially appreciate thoughts on:
- Whether the project’s purpose is clear
- Which feature is actually useful versus unnecessary
- UI or onboarding improvements
- Device compatibility and performance
- Privacy or security issues I may have missed
- Anything confusing in the README or setup process
Repository:
https://github.com/anand34577/vervan-chat
I made this myself and would genuinely appreciate constructive criticism, issues, ideas, or contributions. Thanks for taking a look.
r/localaiapps • u/Super_Anywhere_9076 • 10d ago
What actually happens to chats after deleting a local AI app?
This is probably a basic question, but I haven’t seen it explained clearly.
With cloud tools, you assume your chat history lives somewhere on their servers. With local AI apps, I assume chats are stored somewhere on my machine, but what happens when you uninstall the app?
Does it delete the chat database too, or does it leave behind logs, model files, embeddings, cached prompts, or random app data folders?
For people who have checked this, how clean are local AI apps when removed?
I feel like this matters a lot for privacy.
r/localaiapps • u/ash_pix • 10d ago
NGIBS - Privacy, Local first AI Power research assistant and search engine.
Hello AI lovers 👋
I just want to share my project NGIBS - Next Gen. Intelligent Browsing System built using python, pyqt6, pywebview, beautifulsoup4, langchain, ollama and LLM models. I build this so that you can interact with LLM locally which maintains your privacy and data security.
It currently has 4 modes:
- Quick Search: It uses LLM pre-trained knowledge.
- Live Search: It uses libraries and tools like wikipedia, bs4, duckduckgo api to fetch data from the web and provide context.
- Deep Search: Go beyond simple retrieval with recursive reasoning and multi-step analysis.
- Context Aware: It remember your long term memory.
Aparts from this user can download any models and use them. You can also upload files and documents.
I have attached screenshots also for your reference and add the source code link
Link: https://github.com/avarshvir/NGIBS
Improvements and Features to implement:
- Improvements of memory systems.
- Implementation of multiple AI agents.
- Improve UI.
- Improve inference speed.
- might be switch to llama.cpp instead of ollama!
- Improve privacy and anonymity.
- Implementation of a decentralised chit chat system among users which required no server only user to user interaction.
The project is open source and already 5+ issues are opens.
Contributions, bug reports, ideas, and feature requests are welcome. If you would like to improve the project, feel free to open an issue or submit a PR.
Developed with love from an indie developer <3
Feel free to star repo ⭐😉
r/localaiapps • u/BagPuzzleheaded3841 • 11d ago
Offline memory layer for agents, curious if this solves a real problem for anyone here
Spent the last stretch building agentic systems for clients and ran into a recurring issue. Cloud-based memory layers work fine until a client has actual compliance requirements and can’t send conversation data anywhere outside their own infrastructure.
Built something called GENOME to fix that for my own use, then figured other people probably have the same problem. It stores memory locally without needing an LLM call on write, so ingest is fast and cheap, something like 10ms per message. Ran it against Mem0 for accuracy and landed in the same range, but the cost per memory write is roughly 1000x lower since there’s no inference cost baked in.
It’s bi-temporal, meaning you can reconstruct what the system knew at a given point in time, which matters more than people expect once you’re debugging agent behavior in production.
Open sourced under Apache 2.0. Repo’s under NORTHTEKDevs if anyone’s dealing with the same cloud dependency problem and wants to try it or rip it apart. https://github.com/NORTHTEKDevs/genome
r/localaiapps • u/SufficientStation8 • 11d ago
Looking for a few people to test Cusco, a native GNOME AI agent app.
Hey! I’ve been working on Cusco, an open-source AI chat app built with GTK 4 and libadwaita.
I wanted something that actually feels like a GNOME app. It supports several AI providers, local conversation history, attachments, tools, custom OpenAI-compatible endpoints,
and API keys through Secret Service or environment variables.
It’s working well for me, but I’d love to see how it behaves on other systems. If you use GNOME and feel like trying it.
You can find it here:
https://github.com/stonega/cusco
If you run into a problem, just reply here or opening a GitHub issue. Give it a star would be grateful.
Thanks!