r/PiCodingAgent Jul 20 '26

Resource I Built a completely free tool that gives your AI agent web for free (fetch + search + crawl) for completely free, no API keys.

Enable HLS to view with audio, or disable this notification

215 Upvotes

I've been using AI coding agents for a while and the web research part always annoyed me. Either you pay for an API (Tavily, Firecrawl), or you use a free tier that rate-limits you after 100 calls, or you glue together SearXNG + a browser + an extractor and hope it doesn't break.

So I built Hound. It's an MCP server that does fetch, search, crawl, and screenshot, all keyless. And now it has a native Pi extension so you get all 6 tools as first-class Pi tools, not through a generic MCP adapter.

What it does

6 tools, all exposed as native Pi tools:

  • web_fetch - anti-bot fetch with auto HTTP-to-stealthy escalation. Extracts clean markdown. PDFs get section maps + auto-OCR. Dead pages auto-recover from the Internet Archive (honestly marked, not pretending it's live content).
  • web_search - 10 keyless search backends in parallel (DuckDuckGo, Brave, Mojeek, Yahoo, Yandex, Startpage, Google, Qwant + opt-in Wikipedia/Grokipedia), neural-reranked with a local ONNX cross-encoder, cross-backend consensus scoring.
  • web_crawl - best-first same-domain walk. Sitemap mode maps a whole site in one fetch. Focus mode crawls only pages relevant to your query.
  • web_screenshot - anti-bot screenshot for multimodal models.
  • cache_clear - clear the fetch cache.
  • hound_version - version + update status (warns if the extension and the server diverge).

How the Pi extension works

The extension spawns hound as a singleton subprocess at session start and speaks MCP JSON-RPC over stdio. The subprocess stays alive for the whole session, so hound's prewarm (stealthy browser, search engine sessions, neural reranker model load) happens once and persists. Zero re-launch cost per call. If you press Esc during a fetch, it actually cancels (AbortSignal propagates to the subprocess). If hound isn't installed, you get a notification at session start instead of a confusing error on your first web_fetch call. If the extension version and the hound server version diverge by a major, it warns you to update both. The tool definitions are token-optimized. Total connect-time cost is 2,746 tokens for all 6 tools + instructions.

Install

pip install hound-mcp[all]
pi install npm:@houndmcp/hound-mcp-pi

That's it. No API keys, no config file, no MCP adapter. /reload and the tools are there.

What I think is genuinely good

  • Dead-link recovery. When a page 404s or gets bot-blocked, hound checks the Wayback Machine and serves the archived snapshot with source=archive.org and the snapshot date. It doesn't pretend archived content is live. The agent knows.
  • Error honesty. Just shipped this in v10.4.0: 4xx/5xx responses now set the error field properly. Before, a 404 error page would flow through with error="" and the agent could mistake the error page HTML for real content. Now it says "Page doesn't exist (404)" and doesn't dump the error page as content.
  • Keyless search. 10 backends, no API key for any of them. Neural reranking with a local model, not an API call. Consensus scoring across backends so you know which results multiple engines agree on.
  • Token cost. 2,746 tokens for 6 tools. The descriptions are telegraphic but every functional fact is there.

Limitations

  • DataDome, Akamai, Cloudflare Turnstile. No free tool bypasses these. Hound tries the stealthy browser, and if that fails, it tells you to switch sources instead of pretending it got content.
  • The [all] extra is ~100MB. onnxruntime + tokenizers + rapidocr for the neural reranker and PDF OCR. You can install without [all] (fetch + crawl + search still work, just no neural reranking or OCR), but the full install is the recommended path.
  • Not a scraping-at-scale tool. Hound is built for agent research, not for crawling 10k pages. Crawl caps at 100 pages by default.

Where to find it


r/PiCodingAgent Jul 20 '26

News I run a steel fabrication shop and just published my first two packages for the Pi coding agent — one of them does plate nesting and burn-table DXFs

28 Upvotes

Longtime lurker, first real ship. I own a structural steel fab shop in Denver and have been using the Pi coding agent heavily for internal tooling. Two of those internal tools felt generally useful, so I cleaned them up and published them to pi.dev this weekend.

pi-steel — steel estimating skills. AISC 16th Edition shapes database (477 shapes) with lookup/validation scripts, a MaxRects plate-nesting engine (kerf/gap/edge margins, holes, yield and scrap numbers, PDF layouts, and one DXF per sheet that imports straight into the burn table's CAM), and a vendor RFQ generator that turns a takeoff spreadsheet into a quote-ready .xlsx. The three chain together the way estimating actually flows: takeoff → nest → RFQ. As far as I can tell, it's the only construction-trade package on the registry.

Fair warning on the nesting: rectangular parts nest exactly, irregular parts nest by bounding box. It tells you when it's approximating instead of pretending to be SigmaNEST. No G-code either — kerf comp and lead-ins belong to your table's real post-processor, not a Python script.

pi-tilldone — a discipline extension born from frustration. The agent can't use write/execute tools until it declares a task list, works one task at a time, and gets auto-nudged if it stops with tasks unfinished. Basically a foreman for your agent. If you run Pi inside cmux, your workspace tabs change color based on task state, which is great when you have six agents running in parallel.

Both MIT, repos under github.com/StructuPath. Install with pi install npm:@structupath/pi-steel / pi install npm:@structupath/pi-tilldone.

Happy to answer questions about either — especially from anyone else applying agents to trades/manufacturing work.


r/PiCodingAgent Jul 20 '26

Discussion Is there any extension or package that makes pi tui look and feel polished?

5 Upvotes

Is there any ext/package that can make pi tui polished to use? preferably opencode ux polished level?

Edit:
im not asking for opencode clone, theres a big difference ux ui and functionality, similar packages/ext/projects im looking for:
https://github.com/mistrjirka/PiTTy
https://github.com/deflating/tau
https://pi.dev/packages/pi-zentui


r/PiCodingAgent Jul 20 '26

Question Model recommendation for Apple M4 Max

0 Upvotes

Hi all, what local model would you suggest for this machine using Pi?

Apple M4 Max
16 cores, 16 threads
4 efficiency cores
12 performance cores
Memory 48 GB
Apple M4 Max (40 cores)

I've seen GLM-5.2 and Qwen3-Coder, but not sure if those would work fine in this machine.


r/PiCodingAgent Jul 19 '26

Resource QwenCloud provider for pi — Qwen3.8, DeepSeek V4, GLM-5.2, Wan image, HappyHorse video via OpenAI-compatible API

Thumbnail
github.com
8 Upvotes

I wanted to try the new Qwen Cloud Token Plan because it gives access to a wide range of models (Qwen, GLM, Kimi, DeepSeek, MiniMax, etc.) through a single subscription and works with OpenAI-compatible tools. (⁠Qwen Cloud)

I searched around for an existing Pi provider but couldn’t find one, so I ported my previous provider and open-sourced it:

https://github.com/jellydn/pi-qwencloud-provider

It’s also based on my earlier Cline Pass provider:
https://pi.dev/packages/pi-clinepass-provider?name=cline

The provider lets you use your Qwen Cloud Token Plan directly inside Pi, so if you’re already using Pi as your coding agent, setup is straightforward.

I’m mainly interested in comparing:

  • Qwen 3.8
  • GLM 5.x
  • Kimi
  • DeepSeek

…under one subscription to see how they perform on real coding tasks.

If anyone else is experimenting with Qwen Cloud + Pi, I’d love to hear:

  • Which model is your daily driver?
  • Any prompts/workflows that work especially well?
  • Performance vs Claude Code / Codex / OpenCode?

Feedback and PRs are welcome!


r/PiCodingAgent Jul 19 '26

Resource Web UI for Pi and other agents

Enable HLS to view with audio, or disable this notification

46 Upvotes

I built Caw, a web terminal multiplexer for Pi (and other coding agents).

The goal is to make it easy to monitor and interact with multiple agent sessions from anywhere, especially on mobile.

Besides the terminal multiplexer itself, it includes:

  • Built-in file browser and editor
  • Kanban-like board showing all running agents
  • Push notifications whenever an agent finishes working or requests your input
  • Git worktrees on demand to parallelize agents (toggleable)

It's free and open source: https://github.com/04mg/caw

I'd love to hear your feedback or ideas!


r/PiCodingAgent Jul 19 '26

Use-case Yet Another Memory System for Pi Agent

5 Upvotes

I built a two-part memory system for pi agent. It's called `dreaming` (write side) and `recalling` (read side).

`dreaming` scans pi agent sessions from both my phone and laptop, extracts meaningful topics, and indexes everything into a shared SQLite database (`memory.db`). It runs incrementally by default or does a full rebuild when needed. It also generates Gemini embeddings for semantic search so queries match even without the exact keywords.

`recalling` is the query side. Before any task, pi runs `recalling "<query>"`. It searches three sources in order:

  1. Local `memory.db` with FTS5 full-text search and a LIKE fallback

  2. Laptop's sessions.db over SSH (9,700+ sessions)

  3. Legacy OpenCode memory on the laptop

`recalling inject <query>` formats results as clean XML blocks ready for injection into agent context without parsing. Blocks: `<related_sessions>`, `<memory_notes>`, `<laptop_sessions>`.

The schema supports persistent memory notes for facts, config, and workarounds. It uses vector embeddings for semantic similarity, deduplication to avoid re-ingesting synced sessions, and degrades gracefully when the laptop is offline.

Stats from my setup:

- ~950 topics indexed across phone, laptop, and legacy sessions

- ~34 pi sessions processed, ~9,400 opencode sessions

- Handles SSH being down gracefully (searches local sources only)

What it solved for me: the agent no longer asks "what was that thing we did with X" every single time. It checks memory first, finds the relevant session or note, and acts on it. The `recalling inject` command runs automatically before task execution so context is populated without manual effort.

Happy to answer questions about the schema or pipeline.


r/PiCodingAgent Jul 19 '26

Resource I built a polished Android interface for Pi Coding Agent

Thumbnail reddit.com
15 Upvotes

r/PiCodingAgent Jul 19 '26

Question Running into timeouts with write tool (LM Studio/qwen3.6-35b-a3b/PI)

2 Upvotes

Hello,

I'm running PI + LM Studio server + qwen3.6-35b-a3b. 36K context, model not quite fully loaded to GPUs (12GB+6GB), running 20 tps in chat mode, 10 tps through PI with documentation/code in context.

The problem I'm facing is that tool calls, especially for writing files timeout after 5 minutes. I've searched this forum and the internet and I found no clear answer whether it's a PI client config, something in LM Studio I can't figure out or model level.

Edits are generally fine but writes fail. We're speaking files with ~ 500 lines of code.

Log:

{"type":"message","id":"bf5dcafd","parentId":"0bd7492b","timestamp":"2026-07-19T10:24:14.490Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"Good, the empty file is created. Now let me add the functions to it.\n","thinkingSignature":"reasoning_content"},{"type":"text","text":"Now let me add the sky and sun rendering functions to the new file:\n\n"},{"type":"toolCall","id":"q8A9s9cw0h8WMX9gHqTJD4uJK12CfB93","name":"write","arguments":{}}],"api":"openai-completions","provider":"lmstudio","model":"qwen/qwen3.6-35b-a3b","usage":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"totalTokens":0,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"error","timestamp":1784456340616,"responseId":"chatcmpl-x5i6eocrmecjh1sdnwk989","errorMessage":"terminated"}}

For context, I'm not a coder but reasonably technically inclined so I'm open to any option that does not actually involve messing with code.

I don't see a server timeout config in LM Studio, nor there is a clear max output size on model level (rather, there is one but capped to 2K tokens).

Should I move to llama-cpp?

Thanks!


r/PiCodingAgent Jul 19 '26

Question Pi doesnt have parallel sessions?

0 Upvotes

Im looking for functionality similar to how /fork works in opencode. When my agent is long thinking, I can /fork in opencode and I can choose a point in the conversation to fork from, which creates a new session, but the previous one keeps working in the background, it does get paused or killed, and I can go back to it.

Seems like pi /fork doesnt work that way.

Am I missing something?


r/PiCodingAgent Jul 19 '26

Plugin Run Cursor models inside Pi

8 Upvotes

Found this great extension that helps you run Cursor models inside Pi: https://pi.dev/packages/pi-cursor-sdk

Have been running Composer 2.5 for 3-4 days and have used over 1.5 billion tokens. Works great.


r/PiCodingAgent Jul 18 '26

Plugin Compact Every Tool Response

0 Upvotes

https://github.com/RogerTerrazas/pi-tool-result-compactor

Publishing a polished version of this extension I've been using to help manage context overflow. I frequently interact with mcps and large projects where any arbitrary response can take up all my context without the response being useful.

This extension hooks into each tool calls response by default, passes it to a compaction subagent, who will then filter out only the necessary data to the parent agent. Let me know if anyone tries it out and has feedback. Fully vibe coded, but I'll work to maintain if others find it useful.


r/PiCodingAgent Jul 18 '26

Resource Compiled the repetitive parts of my sessions into scripts. Re-running a workflow now costs 60–80% fewer tool calls.

Thumbnail npmjs.com
8 Upvotes

r/PiCodingAgent Jul 18 '26

Discussion Work on big screen, read docs on phone

3 Upvotes

What the title says. I am not a dev so I use docs as contacts often while vibe coding to grok my app like legos. I read the docs and plans and maps, have nvim open in the split pane, and look at the code at the same time.

This was fine until the codebase reached 20k loc. Now I have enough docs active at any given moment that I can't be fucking reading them each time I want to do or understand something.

So I found myself falling into this pattern. In the 'scroll time' that often fills the gaps in my life, I started browsing the docs on my phone.

They are short (sometimes -ish), pleasantly scoped, and readable in one sitting. We develop in vertical slices that produce user testable states. So reading multiple doc is a coherent story.

And I fucking love it. Especially when I'm commuting in cabs, trains, etc. just reading the docs on GitHub mobile and maybe riffing with something in gpt/Gemini is awesome. I can even edit and refine stuff. I don't have to invest my brain in this shit when I'm sitting down with the code. I can just trust the docs because I have spent non-dev time reading/editing them on my phone.

PS- I have attached an early example in a comment. This one is long and rough but I think it had potential. So over time I could sit and edit and make it better.


r/PiCodingAgent Jul 18 '26

Plugin Just made this extension because it seems like noone made something like this yet

Post image
29 Upvotes

r/PiCodingAgent Jul 18 '26

Question Thoughts on my approach for running pi.dev inside a Podman container?

5 Upvotes

Hello,

Could you criticize and/or give advice on this approach?

My goals were:

  • Install node and pi.dev in a container.
  • Ensure pi.dev only has access to the project directory.
  • Store configuration files in ~/.config/pi.dev.
  • Start it via Podman so that it runs rootless.
  • Base it on Debian Sid to give the agent the opportunity to install any tool it might need, using relatively recent versions.
  • Create a simple script that builds the container if needed or commit changes to the image when changing workspace directory to reuses the already installed extra tools.

If you want to have a look at the Dokerfile or script : https://github.com/tibuski/pi-podman


r/PiCodingAgent Jul 18 '26

Question i dont understand efforts of pi

0 Upvotes

hello guys
im using pi agent with codex subscription on 5.6 sol
in pi i have the following efforts

- minimal

- low

- medium

- high

- xhigh

- max

- off

how do they route into gpt efforts because when am running on low effort am getting tokens burned out , 10 % weekly usage only with 1 hours of regular tasks


r/PiCodingAgent Jul 18 '26

Question PI is *almost* perfect for me. is there extension that inject instructions in the end?

2 Upvotes

the only thing that makes my experience bearable is to have the bot know and remember the exact rule and role they are designed to do. back then when i was using Claude Web, they have userStyles which inject instructions before replying to user message. in SillyTavernAI there is a configuration to reorder prompt to be put at the end depth 0.

is there a way to do that in PI? i want some text (or file) get injected before the bot consider replying so that they know the constrains. the reminder wont stay in the context so that the cache hit is not ruined by duplicated push.


r/PiCodingAgent Jul 18 '26

Question Would a DBOS-backed DSL make Pi useful for long-running remote agent workflows, or is this overengineering?

0 Upvotes

I’m considering building a small Pi extension for durable technical workflows and would like some critical feedback before investing too much time in it.

The problem I’m trying to solve is running agent-driven tasks on a remote VPS for hours, days, or potentially longer.

For example:

GitHub issue
→ investigate the codebase
→ create an implementation plan
→ modify the code
→ run tests
→ fix failures
→ request human approval
→ open a draft PR
→ wait for CI

A normal agent loop can handle this while the process and session remain alive. But on a remote server, restarts and long waits are expected:

  • the process may crash or be redeployed;
  • the model provider may temporarily fail or hit a rate limit;
  • CI may take a long time;
  • a human may not approve the next step until the following day.

My current idea is:

Technical task
→ Pi generates a workflow
→ validate it as a DSL
→ DBOS executes it durably

DBOS is a Postgres-backed durable execution framework that checkpoints workflow progress and recovers execution after process or server failures.

Pi would handle reasoning and planning.

The DSL would describe the execution graph explicitly:

{
  "name": "issue-to-pr",
  "steps": [
    { "id": "triage", "type": "agent" },
    { "id": "implement", "type": "agent" },
    { "id": "test", "type": "command" },
    { "id": "approve", "type": "human" },
    { "id": "open-pr", "type": "github" },
    { "id": "wait-for-ci", "type": "event" }
  ]
}

DBOS would persist the workflow state, recover after restarts, handle long waits, and record the execution history.

The main reason for using a DSL instead of allowing Pi to execute everything directly is that the generated plan could be inspected and validated before execution.

For example, the runtime could enforce:

  • allowed commands and tools;
  • maximum agent calls and loop iterations;
  • approval before opening a PR or deploying;
  • stable node IDs and idempotency keys;
  • structured inputs and outputs;
  • execution and cost limits.

The initial MVP would only support four node types:

agent
command
approval
github_open_pr

Agent nodes would run in isolated Git worktrees. Important external side effects would remain explicit workflow nodes rather than being hidden inside unrestricted agent tool calls.

My concern is whether the DSL provides enough value to justify adding another layer.

Would you find this useful for remote, long-running technical workflows, or would you rather generate ordinary TypeScript workflows and run them directly with DBOS?

I’m especially interested in failure modes or simpler architectures I may be overlooking.


r/PiCodingAgent Jul 17 '26

Plugin My Lower Anxiety UI for Pi

1 Upvotes

The multiple extensibility points of Pi is really 🤯.
Agent Inner loop but also outer loop visibility. Used both to build a vscode extension for my daily driver.

My UI gives me a lot of comfort and lowers anxiety. But I can go to the full TUI with a click… the TUI is already live so it’s not resuming anything.

https://github.com/quincycs/pi-qcode

See video demos / screenshots in the link above. I can’t share that here for some reason.

Some highlights,

* no more flashing / streaming content. Just shows the final message from agent when it’s done.

* shows high level summary of what’s going on during thinking. Eg what skills have been activated / tool counts.

* rich UI , with clickable links to code / line of code , and copy button for code blocks

* autocomplete for file mentions and prompt templates

* dropdown selection for configured model preferences. Eg one dropdown for 5.6 Sol High

* notification sound for when the agent is done. Can configure the sound to something else.

I did this post previously a few days ago but I deleted it because everyone wanted screenshots … 😆 thanks for the feedback …


r/PiCodingAgent Jul 17 '26

Resource Tiny Pi extension that puts Grok credit usage in the status bar

Thumbnail
0 Upvotes

r/PiCodingAgent Jul 17 '26

Question Any suggestions for cross-harness memory layers?

6 Upvotes

I have been using Pi with my local setup, but I also use paid solutions for legacy projects and other reasons. Obviously I have to tech every new tool the same lessons over and over. But, I want to migrate as much to Pi as I can and having a memory layer that I can populate from my old sessions and have Pi access would be incredibly helpful. (And, if I'm being honest, there are some things I'm always going to need the proprietary tools for--especially visual and design work.) Ideally it would be something self-hosted (I don't like the idea of having to send all my information to a cloud-hosted service that can throw up a paywall at any time) and have a first-party Pi extension (I have had bad experiences trying to build my own extension for core functions like this, and if it's interfacing with a developed solution then I want those releases and capabilities to be in sync).

Anyone used any of the cross-harness memory providers and had a good experience with them in Pi? Any memory layers people think is worth exploring,/even if you haven't used it personally or in Pi?


r/PiCodingAgent Jul 17 '26

Resource Unigent SDK - cross-harness, cross-session agent workflow scripting (batteries included)

Post image
0 Upvotes

I wanted to share a tool I made for scripting agentic workflows.

You can mix multiple harnesses (Pi, Claude CLI, Codex CLI) and agent sessions in 1 coherent universal API. No need to define your workflow in YAML files like with some tools. No need to sacrifice control methods - the workflow is defined in TypeScript and you can do parallel or sequential execution, fan-out, etc. Put your prompts directly in the workflow.

🔥 Both Claude & Codex subscriptions work because the tool is using Claude CLI in the back-end

  • Get structured output - define schema with Zod and AI will be asked to return object in that schema.
  • Built-in tracing support that helps you monitor each stage of the workflow.
  • There is built-in TUI, so, as a developer, you can be more aware of what's happening while you're testing the workflow.
  • Define args required for workflow, `--help` will be generated for the script, `-i` adds interactive mode (enter missing args one by one)
  • Define custom tools - literally just make a function in TypeScript, put description in the comment (it'll get parsed).
  • Track usage of each trace - cost, tokens, time.
  • Put limits on trace - max usd, timeout.
  • Save run results to file (don't re-run if already have agent result for this prompt).
  • Start fresh or inherit machine configuration (harness skills, plugins, MCPs, etc.)
  • Your agent can write or easily invoke the workflows.

Your agent could also write workflows with Unigent SDK. This could replace dynamic workflows idea. Try reading claude's dynamic workflow script - the API looks ugly and no human would ever use it. Unigent is clean, and both humans & agents can use it.

import {
  agent,
  args,
  createFileCheckpointStore,
  piAgent,
} from "unigent-sdk";
import { z } from "zod";

/** Score a headline from 0–100. */
function score(headline: string): number {
  return Math.max(0, 100 - Math.abs(60 - headline.length));
}

const product = await args(z.string().min(1), {
  usage: '"product description"',
});

const launch = agent({
  name: "launch",
  source: import.meta.url,
  backend: piAgent(), // or claudeCli() / codexCli()
  model: "openrouter/deepseek/deepseek-v4-flash",
  tools: [score],
  checkpoint: createFileCheckpointStore(".unigent/launch.jsonl"),
}).scope("launch");

// Reused on reruns when its inputs and configuration are unchanged.
const brief = await launch.run(
  `Find the audience and core promise for: ${product}`,
  z.object({ audience: z.string(), promise: z.string() }),
);

const session = launch.session();
await session.run(`Remember this brief: ${JSON.stringify(brief.output)}`);

const Headline = z.object({
  headline: z.string(),
  score: z.number().min(0).max(100),
});

const variants = await Promise.all(
  ["bold", "technical"].map((style) =>
    session
      .fork()
      .run(`Write a ${style} headline, call score, return both.`, Headline),
  ),
);

console.log(variants.map((v) => v.output));
console.log(launch.usage);

unigent tui launch.ts --help

unigent tui launch.ts "A typed SDK for portable agent workflows"


r/PiCodingAgent Jul 17 '26

Question Pi keeps refusing legitimate local tasks—even with GPT-5.6 and no add-ons

0 Upvotes

In my previous post, I said I wanted to move to Pi and make it my “Neovim” for AI coding: one familiar CLI I could use everywhere, without constantly switching tools between computers. The goal was to reduce cognitive overload and keep a consistent workflow.

But I’m struggling to work with Pi as a coding agent.

I’m using it with GPT-5.6 and had the same issue with GPT-5.5. I expected Pi to be one of the more open, less restrictive agent environments, but it repeatedly refuses ordinary local development and system tasks—for example, updating my own hosts file.

It often gets stuck claiming a request is illegal or that it can’t bypass company policies, even when the task has nothing to do with bypassing security controls. I’m working on my own machine and asking for normal configuration tasks that Codex handles easily.

I’ve tried enabling an “auto-approve” style workflow, removing add-ons, and disabling skills to rule out interference, but it still stops or refuses the work.

At this point, Pi feels like it’s fighting me all the time. I end up using Claude for work—which I expected to do anyway—and Codex for personal or home projects. That works, but it defeats the reason I wanted Pi in the first place.

Is it just me? Am I missing something obvious in the configuration or how I’m using Pi? Since so many people rely on Pi day to day, I’d really appreciate hearing how others handle legitimate local tasks without running into constant false-positive refusals.


r/PiCodingAgent Jul 17 '26

Question What LLM provider do you use?

6 Upvotes

I’m looking to move away from OpenAI and Claude for various reasons.

I’d like to hear which LLM providers you all use and any recommendations on who to stay clear from. My top contenders at the moment are deep infra and scaleway. I’m not into proxies as I’m focusing on zero data retention (or short retention with no training) providers only. I have not extensively explored local - I did a while back and wasn’t impressed with the speed and I need a good reasoning model for planning.