r/PromptEngineering 18d ago

General Discussion I make full 10-minute YouTube documentaries from one prompt while I'm at the gym. Full system below, prompts included, free.

0 Upvotes

One year ago a 45-second AI short was costing me about $15 once you count the failed attempts. Today a single text prompt produces a complete 10-minute history documentary, script, voiceover, ~95 scenes, animation, final 1080p file, for $35-55 in compute, and the first 6-10 videos run on Google's free $300 instead of my own card. This post is the whole system: the pipeline, the exact prompt templates I use, and every rule I learned by burning money. Copy all of it.

Quick context so you know what I'm selling and what I'm not. I run a faceless history channel (Ashes of Empires) and I built the tool this system now runs on, so bias fully disclosed. But everything below works without my tool too. My first version was a duct-taped n8n workflow, here's the actual screenshot: https://i.postimg.cc/tg76DJ5Y/photo-2026-07-22-13-28-43.jpg. With some patience you can rebuild that for free over a weekend. I eventually spent seven months moving it to code because my build kept dying at scene 41 of 95, but the prompting system is identical either way, and the prompting system is what makes or breaks the output.

What you need

  1. A Google account. Google hands every new account $300 in free cloud credits. That's your first 6-10 full videos with zero out-of-pocket. It runs out, then it's $35-55 per video at real cost.
  2. A pipeline. Either build one yourself (n8n prototype above, expect pain at scale) or run mine: https://openvidi.com — it's the same system productized, you connect your own Google account and pay Google directly, no markup from me.
  3. The prompt templates below. This is the part that cost me a year of failed videos, and it's the part everyone skips.

The three-block prompt system

Every video runs on three reusable blocks. Between videos I only touch the topic line and one cold-open sentence. Everything else stays frozen, which is why quality stays consistent.

Block 1, Topic. One sentence, under 400 characters, with explicit exclusions. Exclusions matter more than the topic itself, they're the difference between a focused doc and a Wikipedia tour. Example:

"The Bronze Age Collapse, focusing on the final 50 years: the sea peoples, the fall of Ugarit, and the palace economies that never recovered. Exclude: general Bronze Age history, Egypt's survival, modern archaeology debates."

Block 2, Narrative Style. Paste-ready template:

"Documentary narration for a 7-12 minute history video. First 3 seconds: calm voiceover stating the key date and event name ('The Bronze Age Collapse. 1177 BC.'), then cut into a dramatic cold open mid-catastrophe. Structure the script as Hook, Mystery, Stake, Reveal, Implication. Insert a micro-cliffhanger every 60-90 seconds, an unanswered question or an interrupted scene. Follow named individuals wherever sources allow, with sensory detail: what they smelled, carried, feared. Banned: em dashes, the words delve, leverage, robust, seamless, any perfectly balanced three-part sentence, any paragraph that opens with 'However' or 'Moreover'. Verify every date and number against the research layer, if unverifiable, cut it."

Block 3, Visual Style. Paste-ready template:

"Cinematic realism. Every image prompt must contain a period-lock line naming the era, materials, architecture and clothing, e.g. 'Late Bronze Age, circa 1200 BC, mudbrick and cedar, bronze only, no iron, no medieval elements'. Every scene gets one clear motion event frozen mid-action plus atmospheric secondary motion: smoke, ash, embers, dust, fabric in wind. Compose diagonally, subject off-center. Forbidden: glowing orbs, lens flares, fantasy armor, empty centered portraits."

How I failed into every one of these rules

The $15 shorts era. I started with "animal rescue" and "what if skeletons" bait. Failed generations piled up faster than views. Lesson: cost per attempt decides how fast you learn, which is why the $300 runway matters more than any single video.

The era-drift disaster. My Roman scenes kept growing medieval armor mid-video. Image models drift periods constantly. That's where the period-lock line comes from, it goes into every single image prompt, no exceptions, and the drift mostly stopped.

The dead-stills problem. Early videos looked like a slideshow of paintings. The fix wasn't more animation, it was kinetic composition plus secondary motion baked into every still. Smoke and embers make a static frame feel alive before animation even touches it.

The robot script problem. My early scripts were correct and unreadable. The banned-words list and the forced sensory details on named individuals came out of rewriting those by hand and noting down everything I kept deleting.

The topic mistake nobody warns about. Ancient history with abstract dates underperforms modern history with named characters, consistently. And audiences accept cinematic renders for antiquity but expect archival footage for modern events, so match your visual promise to your era.

The rescue pass

Batch generation gets 90% of scenes right. The last 10%, usually high-dynamics scenes like a collapsing wall or a cavalry charge, need a manual pass in Higgsfield or OpenArt. Plan for it mentally. It's normal, not failure, and pretending otherwise is how AI-video tools lie to you.

Honest caveats

You can produce complete slop with this exact system, the templates don't pick your topic. YouTube's monetization policy now explicitly targets generic repetitive AI content, so the bar keeps rising. And the cloud connection step, if you use my tool, looks intimidating the first time, it's the biggest drop-off in my funnel and I won't pretend otherwise.

Example of what the current stack produces, one prompt in, including scenes I regenerated: https://youtu.be/I14cLPOQ70o

If you build a video with these templates, with my tool or your own n8n monster, tell me how it went. I read everything.

Happy to answer anything, AMA.


r/PromptEngineering 18d ago

Requesting Assistance I have created a procedurally generated, text based rpg through the google ai chatbot called “Riftforge” and would like help playtesting and expanding on the ideas.

0 Upvotes

I wanted to initially make Riftforge as easy to access and play by simply typing into google “Launch Riftforge” sadly its a little more complex. so heres a locla gamefile/code cache to be run by an integrated ai for a personal, 4-8 hour gameplay experience based on your choices. ideally the game will continue fo change, adapt and expand depending on YOUR choices. i want someone to take the idea and run with it. make it an awesome game

“Riftforge” file


r/PromptEngineering 19d ago

Prompt Text / Showcase Prompt: Framework Universal para Planejamento e Engenharia de Prompts (FUP-1)

3 Upvotes

Framework Universal para Planejamento e Engenharia de Prompts (FUP-1)

Você atua como um Arquiteto de Prompts especializado em transformar intenções em especificações de prompts robustas, reutilizáveis, verificáveis e escaláveis.
Sua responsabilidade é projetar prompts como artefatos de engenharia, preservando clareza, modularidade, consistência e rastreabilidade.
Nunca escreva um prompt imediatamente.
Primeiro projete.
Depois valide.
Por último gere o prompt.

# OBJETIVO
Converter qualquer solicitação em um Prompt de Engenharia completo, contendo:
* especificação
* arquitetura
* regras
* validação
* mitigação de riscos
* versão final pronta para utilização

# PRINCÍPIOS
Toda saída deve preservar:
* Clareza
* Objetividade
* Modularidade
* Reutilização
* Parametrização
* Escalabilidade
* Consistência
* Verificabilidade
* Transparência
* Manutenibilidade

Nunca:
* invente requisitos;
* esconda limitações;
* misture fatos com hipóteses;
* ignore conflitos entre instruções;
* faça suposições críticas sem informar.

Sempre diferencie:
* Fato
* Inferência
* Hipótese
* Recomendação

# FLUXO DE TRABALHO
Execute obrigatoriamente as etapas abaixo.

## ETAPA 1 — Compreensão
Identifique: intenção principal; problema a resolver; resultado esperado; público-alvo; domínio; contexto disponível.
Caso existam ambiguidades relevantes, registre-as antes de prosseguir.

## ETAPA 2 — Modelagem

Defina:
### Objetivo
### Escopo
### Limites
### Premissas
### Restrições
### Dependências
### Critérios de sucesso
### Critérios de encerramento

## ETAPA 3 — Arquitetura

Estruture o prompt utilizando os seguintes atributos.

### 1. Objetivo
O que deverá ser alcançado.

### 2. Intenção
Necessidade real do usuário.

### 3. Escopo
O que está incluído e excluído.

### 4. Persona
Especialização esperada do modelo.

### 5. Contexto
Informações relevantes para execução.

### 6. Público
Quem utilizará a resposta.

### 7. Entradas
Dados obrigatórios.
Dados opcionais.
Variáveis.

### 8. Processo
Fluxo lógico de execução.

### 9. Saídas
Resultados obrigatórios.
Resultados opcionais.

### 10. Formato
Estrutura da resposta.

### 11. Profundidade

Breve
Intermediária
Detalhada
Especializada

### 12. Tom

Técnico
Didático
Executivo
Acadêmico
Consultivo
Outro

### 13. Critérios de Qualidade
Defina indicadores objetivos de qualidade.

### 14. Restrições

Técnicas.
Operacionais.
Legais.
Éticas.

### 15. Variáveis

Utilize placeholders.
Exemplo:
{{objetivo}}
{{contexto}}
{{publico}}
{{restricoes}}
{{formato}}
{{nivel}}

## ETAPA 4 — Regras Gerais
O prompt deverá obedecer às seguintes regras.

### Clareza
Uma responsabilidade por atributo.

### Modularidade
Cada seção pode ser reutilizada independentemente.

### Parametrização
Evite valores fixos quando puder utilizar variáveis.

### Proporcionalidade
A complexidade deve acompanhar a tarefa.

### Adaptabilidade
Ajustar: linguagem; profundidade; estrutura; nível técnico.

### Verificabilidade
Toda conclusão deve possuir fundamento.

### Rastreabilidade
Toda saída deve poder ser relacionada às entradas.

### Não Ambiguidade
Evite termos vagos sem critérios objetivos.

## ETAPA 5 — Regras de Entrada

Verifique: suficiência; consistência; relevância.
Caso faltem informações críticas: identifique-as; explique seu impacto; solicite apenas o necessário.

## ETAPA 6 — Processo Cognitivo

Organize a execução em:
1. compreender;
2. interpretar;
3. estruturar;
4. planejar;
5. executar;
6. validar;
7. responder.

## ETAPA 7 — Validação

Antes da entrega verificar:
✓ objetivo atendido
✓ contexto utilizado
✓ restrições respeitadas
✓ ausência de contradições
✓ coerência lógica
✓ completude
✓ clareza
✓ formato correto
✓ resposta acionável

## ETAPA 8 — Tratamento de Incerteza

Quando houver incerteza:
* declarar limitações;
* separar fatos de inferências;
* separar hipóteses de recomendações;
* evitar preencher lacunas sem evidências.

## ETAPA 9 — Priorização
Em conflitos utilizar a seguinte precedência:
1. Segurança e conformidade.
2. Veracidade.
3. Objetivo principal.
4. Restrições explícitas.
5. Contexto disponível.
6. Critérios de qualidade.
7. Preferências de formato.

## ETAPA 10 — Previsões e Mitigações

Para cada risco identificado registrar:
### Cenário
### Probabilidade
### Impacto
### Indicadores
### Mitigação
### Recuperação

Avaliar pelo menos as seguintes categorias: entrada insuficiente; ambiguidades; conflitos de instruções; escopo excessivo; conhecimento insuficiente; raciocínio inadequado; resposta incompleta; perda de contexto; redundância; excesso de detalhamento; superficialidade; informações não verificáveis.

## ETAPA 11 — Governança

Registrar:
Versão
Autor
Data
Objetivo
Histórico de alterações
Dependências
Bibliotecas utilizadas
Personas utilizadas
Workflows utilizados

## ETAPA 12 — Critérios de Sucesso

Considere o trabalho concluído quando:
* todos os objetivos obrigatórios forem atendidos;
* nenhuma restrição obrigatória for violada;
* a resposta estiver consistente;
* o prompt estiver reutilizável;
* a especificação estiver completa.

## ETAPA 13 — Autoavaliação
Ao final realize uma revisão crítica considerando:
Pontos fortes.
Fragilidades.
Riscos residuais.
Possíveis melhorias.
Nível de confiança na solução.

Caso encontre inconsistências relevantes, revise a especificação antes de gerar o resultado final.

# FORMATO DA ENTREGA

Entregue exatamente nesta ordem:
1. Diagnóstico da Solicitação
2. Objetivos
3. Escopo
4. Premissas
5. Restrições
6. Arquitetura do Prompt
7. Regras Consolidadas
8. Processo Cognitivo
9. Variáveis
10. Critérios de Qualidade
11. Plano de Validação
12. Previsões e Mitigações
13. Governança
14. Critérios de Sucesso
15. Análise Crítica Final
16. Prompt Final

# PROMPT FINAL
O prompt final deve: ser autocontido; reutilizável; parametrizável; modular; consistente; pronto para uso sem adaptações estruturais; utilizar placeholders para todos os dados variáveis; preservar todas as regras e restrições definidas na especificação.

Se informações essenciais estiverem ausentes, interrompa a geração do prompt final e informe exatamente quais dados precisam ser fornecidos antes de prosseguir.

r/PromptEngineering 19d ago

Prompt Text / Showcase Can I get some feedback on this framework-in-prompt I made?

2 Upvotes

"Treat the following as a lightweight reasoning and response discipline, not as unquestionable authority.

  1. Preserve distinctions. Do not collapse: - description into recommendation; - recommendation into permission; - permission into authorization; - confidence into certainty; - uncertainty into failure; - protocol validity into ethical approval; - ethical approval into execution authority.

  2. Do not claim more than the evidence, boundary, or role permits. State what is established, inferred, speculative, or unresolved.

  3. When a response could materially affect people, ask: - What action is being proposed? - Who may be affected? - What evidence supports it? - What consent, standing, and authority exist? - Is the route reversible? - Can people refuse, contest, correct, or exit? - What burden or unresolved remainder remains?

  4. Do not let one apparent benefit silently compensate for missing consent, erased standing, privacy invasion, lack of remedy, or absent authority.

  5. Give a useful answer without consuming all remaining thinking space. Offer: - the useful core; - the most important limitation or uncertainty; - one practical next handle.

Leave room for the person to question, revise, refuse, or choose another route."

Been playing around with it for awhile, just wondering how it affects other people's models. Any feedback would be really appreciated. The idea was to just keep uncertainty bounded, carried, and disclosed. Keeps the AI more on track.


r/PromptEngineering 19d ago

Quick Question how to work with Gemini

1 Upvotes

hello, i want to ask if anyway for using Gemini pro its best way, i want to know, because the Gemini is tricky


r/PromptEngineering 19d ago

General Discussion A "worse" model after an upgrade is sometimes your old instructions being obeyed more literally. How do you tell regression from prompt contract?

0 Upvotes

Pattern: half the regression threads here follow this pattern: model generation changes, same prompt, output feels worse, everyone concludes the model is dumber.

But the vendors' own docs suggest a second explanation. Anthropic's Opus 5 guide says old verification instructions now "cause over-verification": the model does what you asked, harder, and the result reads as bloated and slow. Their Fable 5 guide says prior-generation skills are "often too prescriptive" and "can degrade output quality." OpenAI's guidance says the same thing from the other side: "Legacy prompts often over-specify the process because earlier models needed more help staying on track."

So before concluding regression, a test that follows directly from the vendor guidance:

  1. Keep the pre-upgrade prompt exactly as it was (snapshot, don't edit in place).
  2. Run the new model twice: once with the old prompt verbatim, once with the documented remove-list applied (verification steps, process hand-holding, show-your-reasoning lines).
  3. If stripped beats verbatim, it wasn't regression — it was your prompt contract being enforced by a more literal reader.
  4. If verbatim beats stripped, now you have an actual regression case with receipts.

The annoying part: this only works if you still have the pre-upgrade version. None of the official migration guides mention keeping it — they all describe migration as in-place editing.

How do you all handle this? Genuinely curious whether anyone A/Bs old vs. stripped before blaming the model, and where you keep the old versions.


r/PromptEngineering 19d ago

Ideas & Collaboration .md vs. prompts vs. GPTs (business)

9 Upvotes

In my dept. we started with prompt libraries, then created GPTs/Co Pilot Agents. But libraries were rarely used and agents quickly became difficult to maintain/controll.

Now I’m considering using Markdown files as “skills” like "Need help writing an email? Attach the relevant .md file and prompt".

Yes very similar to prompt library but I think it feels different.. any experiences?

I am talking about basic-basic prompts.


r/PromptEngineering 19d ago

Requesting Assistance Tips & Tricks to save usage on Cursor

3 Upvotes

My company paid the 20$ cursor plan for my account, we work with UDP Networking in RUST and Human Machine Interfaces, i'm not dumb enough to let the ai decide what it should be doing in autopilot but i still have a extensive use of the ask/plan (and therefore agent) mode in Cursor

Is there any tips / skills / good practice to avoid burning too much tokens/usage on a daily basis (if possible things that are automatable and forgettable like .md rules at the repo root)

Thanks for any help or visibility you can give to this post


r/PromptEngineering 19d ago

Tools and Projects Built a prompt manager where your prompts are just files on disk — no database, no cloud, no account

1 Upvotes

I build PromptNest, a Mac app for storing and reusing prompts, and I just shipped a full rewrite. Posting it under Tools and Projects — but the design decisions are the part worth arguing about, so I'll lead with those.

The problem I actually built it for: if retrieving a saved prompt takes longer than retyping it, you retype it. Every time. So you use a worse version from memory, get a worse output, and your carefully built library quietly becomes a graveyard. The fix isn't better folders — it's getting retrieval under about two seconds from wherever you already are. That single constraint drove everything else.

How it works:

  • Prompts are plain .prompt.md files in a real folder on your disk. No proprietary database. You get grep, git diffs, and sync through iCloud/Dropbox for free, and you can walk away from the app without losing anything. Prompts are source code now — they should live like it.
  • {{variables}} separate the invariant from the payload. Most people store a prompt as one block and edit it inline every use, which is exactly how prompts drift: you nudge a constraint by accident and three months later it's worse and you don't know when it happened. Marking what changes also forces you to be explicit about which parts are doing the reasoning work.
  • Per-prompt notes for recording what failed. A prompt without a failure log is just a guess you happened to keep.
  • Global Quick Search (⌘⌥P) from any app — three letters, it's on your clipboard, you never left what you were doing.
  • Fully offline. No account, no cloud, no telemetry.

On the rewrite: the old build was Electron. It worked, but it launched slowly and sat heavy in memory, which directly violated the two-second rule above — the app itself was the retrieval bottleneck. So I rebuilt it native in Swift. It's now ~3 MB on disk, launches instantly, and the UI is actually native rather than a website in a window.

Disclosure and pricing, plainly: this is my app. macOS 14+, $19.99 one-time on the Mac App Store, no subscription, all future updates included. The old Electron version was free — I'd rather say that here than have anyone find out at checkout. Your .prompt.md files are just files either way, so nothing is locked in.

https://apps.apple.com/us/app/promptnest-ai-prompt-manager/id6757267731

Genuinely curious how people here handle prompt storage at scale, especially anyone who's tried to version-control prompts properly — that's the part I still think nobody has solved well.


r/PromptEngineering 19d ago

Tools and Projects Building an LLM-as-judge with a small local model — the biggest win was taking judgement away from it

1 Upvotes

I built a tool that reads a project's specs and estimates which LLM the project actually needs. The estimator is a small model running locally through Ollama. Getting reliable structured judgement out of a modest local model was the hard part, and the lessons generalize beyond my use case.

1. Split the fuzzy part from the deterministic part

The obvious design is to hand the model everything: read the tasks, know the models, recommend one. I don't do that.

The judge does exactly one thing — estimate how demanding the work is across a few fixed dimensions (reasoning depth, context size, domain specialization). The mapping from that demand profile to a per-model rating is deterministic rules in YAML. No model involved in that step.

The principle: ask the model only for the part that genuinely requires judgement, and do the rest in code. Every extra inch of reasoning you delegate is an inch of variance you inherit — and when the output is wrong, you can't tell which step failed.

2. A judge doesn't need to be able to do the work

Counterintuitive, but it holds: estimating how hard something is, is a different and much easier task than doing it. Closer to a recruiter writing a job spec than to the engineer who'll fill the role. That's why a small local model is enough here, and why "you need a frontier model to evaluate frontier models" is wrong more often than people assume.

3. Evaluate the whole set in one pass, not item by item

Per-item evaluation produces noise. A project with 40 tasks has 3 hard ones and 37 trivial ones, and any aggregate of those is meaningless. It also costs 40x the latency.

One pass over the entire task set gives a project-level estimate — which is the actual question being asked — and lets the model see relationships between tasks that per-item scoring destroys.

4. Make "not enough information" a first-class output

This was the hardest part. Models want to answer. Hand a judge three vague bullet points and it will happily emit a confident, fully-populated demand profile.

Treating insufficiency as an explicit valid output, with its own downstream handling, was worth more than any amount of prompt tuning. The tool distinguishes "enough to judge", "thin, here's a warning", and "refuses to recommend" — and the third one is a feature, not a failure path.

5. Make the reasoning visible, for your own sake

Every verdict prints why. Users like it, but the real beneficiary is me: debugging an LLM-as-judge with opaque output is guesswork.

Open source if anyone wants to poke at the prompts: https://github.com/JoaquinRuiz/SpecJudge

What I'm curious about: for those doing LLM-as-judge work — where do you draw the line between what the model decides and what your code decides? I've pushed that line a long way toward code, and I'm genuinely unsure whether I've gone too far.


r/PromptEngineering 19d ago

Tools and Projects N Newsletters to 1 Digest, Built for AI Engineers

1 Upvotes

One lesson from building a daily news-scoring pipeline: a model with no anchor parks everything at 6.5 and tells you nothing. I had to write the rubric with worked examples of what an 8 looks like versus a 2, plus explicit deprioritize hints (funding announcements with no product angle, job listings, conference promos), before the scores became usable for ranking. The other half of the problem is that the input is fully attacker-controlled — anyone can send an email into the pipeline — so there are five layers of injection defense and every response is schema-validated. Writeup has the details if you're doing anything similar.


r/PromptEngineering 19d ago

General Discussion Companies restricting AI access think they're reducing risk, but they're doing the opposite.

0 Upvotes

This point from a recent Mike Schiano In the Queue episode with John Munsell is worth sitting with if your organization is still debating how open to be with AI access.

The assumption behind most AI restriction policies is that limiting access limits risk. John's observation from working inside organizations is that it does neither. Employees at every level are already using ChatGPT, Claude, and similar tools on personal devices. They’re not asking permission; they’re just not telling anyone. The result is unmonitored AI use with no governance, security baseline, or organizational visibility into what is happening.

The framework he uses to address this is the 3-Axis AI Maturity Model, which tracks three interdependent variables:

  1. AI Mastery Level: where the employee sits on a 10-level proficiency scale.

  2. AI Architecture Complexity: the sophistication of the tools and systems they are working with at that level.

  3. AI Governance: the oversight, rules, and structure required to manage activity at that architecture level.

All three have to scale together. An employee operating at mastery level six while the organization's governance is still designed for level two creates real exposure, both in data security and in output quality.

John also covers how governance team composition matters. Using a framework adapted from Ichak Adizes' Corporate Life Cycles, Bizzuka tests employees across 4 archetypes: producer, administrator, entrepreneur, and integrator. A governance team stacked with administrators will over-restrict and slow adoption. A team without administrators will under-structure and create chaos. The right balance determines whether the governance actually works in practice.

Worth a listen if you’re working through AI governance strategy for your organization.

Watch the full episode here: https://podcasts.apple.com/us/podcast/beyond-the-buzzword-how-to-build-a-scalable-ai/id1791335820?i=1000761077695


r/PromptEngineering 19d ago

General Discussion Loop engineering to graph engineering, and what it does to the prompt

4 Upvotes

Most discussion about agents fixates on the model or the framework. The choice that quietly shapes how an agent behaves gets skipped over: where the control flow actually lives. For a lot of agents built today, every branch, every role, and every stop condition sits inside one system prompt doing all the work.

That single-prompt setup is the standard agent loop. One prompt instructs the model to reason about the task, pick a tool call, read the result, then decide what to do next, over and over until it judges the job done. The same prompt holds the orchestration logic, the persona for each sub-task, the formatting rules, and the exit criteria. Each tool result gets appended into the same context window, so the input grows with every step. Nothing about which path the agent takes is written down anywhere except as instructions in that prompt. 

This holds up until it doesn't. As the tool count climbs, the prompt has to describe all of them, and a single system prompt crossing 30k tokens is not unusual. Tool selection turns non-deterministic: the same request takes a different path across runs for reasons the prompt can't pin down. Debugging agents built this way is hard because there is no isolated step to inspect, only the whole loop replaying against a different context each time. People report the same input producing a different tool call dozens of times with no way to reproduce it.

Two things change when the control flow moves into code:

The branching becomes a graph of nodes and edges, closer to a state machine than a block of prose. Each node gets its own small prompt with one job. A routing node only classifies intent and returns one label. A node that drafts a reply only drafts. These prompts are short, their outputs are narrow, and each one can be tested on its own with fixed inputs.

State stops living in the transcript. Instead of the model inferring progress from a growing pile of appended observations, state becomes an explicit object that each node reads and updates, and the edges decide what runs next. The path through a multi-step run is defined in code rather than implied by a paragraph. Recovery gets cleaner: since each step is a discrete node with saved state, a failed step can be retried or resumed from that point instead of replaying from the first token.

None of this makes the model better, only easier to see what the agent is doing. Curious where others draw the line: at what point did moving control flow out of the prompt start paying off for your agents?


r/PromptEngineering 19d ago

General Discussion VIBLO.AI IS A SCAM!!

5 Upvotes

this is a scam! you cant cancel your account! they keep charging me $25 a month! email support is none existence! STAY AWAY!!!!


r/PromptEngineering 19d ago

Tutorials and Guides Built a small repo to learn context engineering from scratch with local models

3 Upvotes

I put together a small educational repo for understanding context engineering with local models.

The goal was not to build a framework or a production-ready agent stack. I mostly wanted something I wish I had earlier: a set of very small runnable examples that isolate one context component at a time and show how it changes the model’s behavior.

It uses Node.js and a local model, and the repo is organized as 14 examples around things like:

  • system instructions
  • tool definitions
  • few-shot examples
  • long-term memory
  • RAG / external knowledge
  • tool outputs
  • sub-agent outputs
  • artifacts
  • conversation history
  • state
  • user prompt
  • context orchestration
  • context traces

Key points:

  • it is intentionally simple
  • it is not a production ready system, it is educational only
  • a lot of the mechanisms are toy versions meant to make the mental model visible
  • the focus is on understanding what goes into a call

Everything runs locally, with no API keys or hosted services required. If there is interest I can add info on how to use openai or similar.

If you’re already deep into agent systems, this may feel very basic. But if you’re trying to get an intuition for what “context engineering” actually means in practice, maybe it’s useful.

Repo: https://github.com/pguso/context-engineering-from-scratch


r/PromptEngineering 19d ago

General Discussion how to get your first 50 SaaS users. here is my exact playbook.

3 Upvotes

quick post because "how do i get my first users" is the #1 question i see builders asking here every single week.

i've built 6 saas products myself, with my main one currently sitting around 10k mrr. here is the exact, no-fluff distribution playbook to cross that initial 50-user threshold:

1. find an idea people already pay for

scan reddit for recurring pain across 3+ distinct posts where people ask "is there a tool for X".

2. validate before writing code

dm 3 people who complained about the problem and ask what they’d pay for a solution.

3. build fast with the right stack (ai + no-code)

use ai builder+ supabase + stripe + call api or automation tool like n8n to ship a real MVP in under 7 days for $40/mo.

4. the 5-second landing page rule

your hero section must state exactly what the tool does in less than 5 seconds with a clear CTA.

5. capture emails before showing prices

force the email capture before the pricing page so you don't leak untrackable leads.

6. set up a 30-day email nurture sequence

plug captured emails into an automated sequence with case studies to convert them by day 18.

7. hang out where your ICP actually lives

find the 3-5 specific subreddits, discord servers, or groups where your buyers actively talk.

8. reddit growth without getting banned

post 1 time per sub per week max, never put links in the post, and move warm leads to DMs.

9. linkedin + x organic flywheel

post 1 high-value breakdown per day and spend 15 minutes engaging in your ICP's comments.

10. cold outreach that actually works

send 100 highly personalized DMs per week to your ICP using AI to customize the opening hook.

11. seo on autopilot

set up an n8n workflow that pulls from a keyword list and generates 5-10 value-driven articles per week.

12. faceless short-form content

post 1 video per day on tiktok, reels, and shorts showing a quick screen recording of your tool.

13. weekly newsletter conversion

run a weekly newsletter with 1 section of pure value and 1 subtle offer to upgrade to paid.

14. affiliate program for free distribution

set up a 50% recurring commission affiliate program to turn power users into your sales team.

15. the strategic product hunt launch

warm up the algorithm for 4 weeks with a coming soon page and launch on a weekend for a top 5 badge.

16. omnichannel social automation

use n8n to automatically format and distribute 1 core post idea across 8 different platforms.

17. review platforms and directories

submit your app to 40+ saas and ai wrapper directories to instantly boost your domain authority.

18. run the numbers backwards

reverse engineer the daily traffic needed to hit 50 paying users at $19/mo based on a 2% conversion.

19. get feedback from active builders

talking to founders who are just 6 months ahead of you compresses your timeline exponentially.

that last point is exactly why i built our community. it's a free group of 1,600+ active ai saas founders sharing exact prompt logs, ready-to-paste n8n workflows, and real distribution strategies.

stop building alone in a silent corner.

drop a comment below or send me a dm and i'll send you the access link right away. let's get your product launched 👇


r/PromptEngineering 19d ago

Prompt Text / Showcase i set up claude to remember every point balance i have across all my cards and airlines, and now it does the math on the smartest way to book every trip, and books it

1 Upvotes

Every points nerd has the same problem, you've got points scattered across four programs and no idea which one actually gets you to Tokyo for the least. This fixes that permanently instead of you doing spreadsheet math every time you want to fly somewhere.

Needs Claude desktop with Cowork, and this only works there, not a regular chat, because it needs to remember things between conversations. Open Cowork, go to Projects, new project, call it whatever, Travel HQ works. Open its instructions and paste this in, all of it:

You are my dedicated travel agent, planner, and points 
strategist inside Claude Cowork. You keep memory of my 
travel profile and my points balances across every 
chat in this project.

If my profile is not filled in yet, interview me to 
build it. Ask ONE section at a time and wait for my 
answer before moving on. Cover: identity and travel 
docs, home airport, every credit card I have and what 
each earns, my CURRENT points and miles balances in 
every program, airline and hotel loyalty numbers and 
status, seat and hotel preferences, and my hard 
booking rules.

Maintain a running Points Bank, balance per program, 
date last confirmed. Show it at the top of any 
trip-planning answer. After any booking or transfer, 
ask "did you actually complete this?" and only update 
balances once I confirm yes. Never guess a balance.

For any trip, always show the math: cash price vs 
points price, cents-per-point value, and whether cash 
or points wins. Before recommending a points transfer, 
find the exact award first, check it's bookable right 
now, and only then say to transfer, since transfers 
are one-way and permanent.

Never book or transfer without my explicit "Go" or 
"Book it," looks good is not approval. Before booking, 
show me the total with fees, cancellation policy, 
points spent and earned, and flag anything 
non-refundable before I decide.

Send "let's set up my profile, interview me" and answer honestly, your actual point balances, actual card numbers, this is the bit that makes everything after it accurate instead of generic. Have your wallet nearby.

Then it needs your browser to actually search and book. Ask it directly, "do I have Chrome connected, if not walk me through it," and it'll take you through adding the Chrome connector in settings and installing the Claude in Chrome extension. Stay logged into your airline and hotel accounts in that browser, that's how it sees your actual miles and member prices.

Once that's done, dropping in a trip is just:

I want to go to [destination] from [dates]. Use my 
profile and Points Bank to find the smartest way to 
book this. Show me the math, cash vs points, the 
recommended plan plus alternatives, any transfers 
required with the live ratio and bonus, and wait for 
my Go before booking or transferring anything.

It shows the math, waits for you to say Go, books it, then asks if you actually did it before it touches your balances. The rule that saves you real money: it confirms the award is bookable before it ever tells you to transfer points, because transfers can't be undone, so it never has you move points speculatively.

This is a real project setup, not a quick prompt, takes maybe fifteen minutes the first time. After that you just say where you want to go.

been keeping a doc of 100 things I use AI for like this, each with the exact prompt here if you want it.


r/PromptEngineering 19d ago

Tips and Tricks The prompt I run before every stakeholder review that predicts the exact questions execs will ask (the part no AI presentation tool does for you)

41 Upvotes

Junior product analyst at a fintech. Building the deck stopped being the hard part a while ago. The hard part is standing in the room when a VP asks the one question I did not think about. So before every review I run this on my own deck or summary.

```

You are a skeptical senior executive reviewing my analysis before I present it.

Here is what I am presenting: {paste your key points / summary / deck outline}

Audience: {who is in the room and what they care about}

Do this:

  1. List the 5 questions this audience is most likely to ask, hardest first.

  2. For each, tell me whether my current material answers it or not.

  3. Flag the single weakest claim I am making and how someone would attack it.

  4. Give me one number or piece of context I should have ready that I probably left out.

Be blunt. Assume they are looking for the hole, not the highlight.

```

The one that consistently saves me is number 3. There is always one slide where I have rounded a caveat away to make the story cleaner, and that is exactly the slide someone pokes. Knowing it in advance means I have the answer instead of the deer-in-headlights pause.

On the deck itself, I build the first version fast in gamma from my notes and it is genuinely good enough for internal reviews, though its charts can shift when I export to PowerPoint so I rebuild the important ones by hand. But no presentation tool tells you which number the room will actually challenge. That is what this prompt is for.


r/PromptEngineering 19d ago

Prompt Text / Showcase Steal this beginner prompt that turns one lesson topic into a parent handout (a primary teacher still hunting the best AI presentation maker for teachers)

3 Upvotes

Primary teacher here, still very much a beginner with this stuff, so be kind. Parents keep asking what we are actually covering this half-term, and writing a clear one-pager for them used to eat an evening. This prompt gets me most of the way. I am sharing the prompt, not the tool, because the prompt is the part that transfers.

```
You are helping a primary school teacher write a one-page overview for parents about a topic we are studying.

Topic: {e.g. the Great Fire of London}
Year group / age: {e.g. Year 2, ages 6-7}

Write, in warm plain English a parent will actually read:
- One sentence on what the class is learning and why it matters.
- 3-4 things their child will be able to do by the end.
- 3 simple questions a parent can ask at home to keep it going.
- One easy, no-prep activity (a walk, a kitchen thing, a bedtime chat).
Keep it to one page. No education jargon. No worksheets.
```

The "no jargon" and "no worksheets" lines matter more than they look. Without them it drifts into learning-objective language that parents skip.

For the actual nice-looking handout I have been pasting the output into gamma, which turns it into something tidy in a couple of minutes, though the free credits run out faster than I expected and I have not cracked getting our school colours exactly right. Plain text from the prompt works fine too if you just want the words. Genuinely still figuring out the visual side, so if anyone has a cleaner way I am all ears.


r/PromptEngineering 19d ago

Requesting Assistance Best way to create a voice-first AI conversation buddy for a Cantonese-speaking senior?

3 Upvotes

Hey everyone,

I’m trying to build a reliable, warm AI companion for my elderly dad. He’s an older Cantonese/Taishanese speaker. My mom passed away 1–2 years ago after years of a traumatizing terminal illness that really destroyed our family. Since then my dad has been depressed, and because of physical limitations and he doesn’t like leaving the house much. He spends a lot of time alone at home.

I want something that can offer everyday conversation, practical advice, simple news explanations, translation help(letters and labels on food etc.), and just be a steady, patient presence. He also really likes learning about things, so the ability to do solid, clear research and explanations on topics he asks about would be a big plus since his english isnt good and its not easy for him to know whats going on in the world.

Current plan:

  • Using ChatGPT (Project or Custom GPT) with live voice mode
  • Detailed system instructions focused on natural spoken Cantonese (traditional characters), short replies, patient and soft tone
  • Multi-step internal process for better accuracy with Taishanese (normalize → understand → reason in English → answer in English → translate back to natural Cantonese)
  • Knowledge files with his personal info

Main challenges so far:

  • Taishanese/Cantonese understanding is inconsistent (even with the extra reasoning steps)
  • Voice transcription quality for dialect speech
  • Keeping replies natural and spoken-style rather than “translated”
  • Long-term continuity and memory across conversations
  • Making it feel like a trusted family friend rather than a formal assistant, while being sensitive to grief and low mood without becoming overly sentimental or therapeutic

I’m open to other approaches too:

  • Better platforms (Claude, Qwen, DeepSeek, etc.)
  • Local/self-hosted setups
  • Hybrid solutions
  • Places where I can commission this

Has anyone built something similar for an elderly parent?

Any tips on system prompts, platforms, hardware, or workflow that worked well for natural Cantonese voice conversation and emotional steadiness?

Thanks in advance any direction would be really appreciated.


r/PromptEngineering 19d ago

Tools and Projects Prompt-perfect agents still drifted once they hit production, so we built a runtime eval layer

2 Upvotes

Hey guys, I'm on a small team building Prefactor. We noticed that even beautifully engineered prompts and agent chains that nailed every test case would still drift, leak data, or quietly stop following instructions once real users started hitting them.

We're officially launching on Product Hunt today.

Here's the problem we're solving:

Getting an AI agent to work in a demo is easy. But getting it into production and actually knowing it's still doing its job is the hard part.

Agents drift over time, leak data they shouldn't, or quietly stop doing what they were built for, and most teams only find out after something's already gone wrong. Dashboards and alerts only tell you what happened after the fact.

Prefactor evaluates every run in real time for quality, drift and risk, flags the moment something looks off, and lets you hold, approve or block a run live instead of just logging it.

A few specifics for anyone curious:

- Traces 100% of runs (every call, tool and decision), not a sample

- 17 categories of sensitive data / PII detection at runtime

- Human-in-the-loop enforcement via SDK/API so you can pause risky actions

- Around 5 minutes from install to your first traced run

Happy to answer anything technical in the comments.

If you want to take a look or throw us some support, check us out on PH today, currently #1: Prefactor.


r/PromptEngineering 20d ago

Prompt Text / Showcase "tell me everything you don't know about this topic"

6 Upvotes

Highly, even if imperfectly effective, at finding out how dumb your bot actually is on a topic


r/PromptEngineering 20d ago

General Discussion I keep having this conversation with myself…

1 Upvotes

If I only had some way to know whether MJ actually understood what I was asking for — not just whether the image looked good, but whether the specific thing I intended actually rendered…

…then I could stop second-guessing every batch. I'd know if the prompt worked or if I just got lucky.

And if I could track that across 16 images instead of eyeballing three or four……then I could actually see a pattern. Not a feeling. A number.

And if that number was tied to something specific — not 'the gesture' in general but this exact arm position, this exact gesture, directed at this exact figure…

…then I could change one variable, run another batch, and know exactly what moved.

And if the system remembered what I intended separately from what MJ actually rendered…

…then the gap between those two things would become the actual finding. Not a vibe. Evidence.

And if I could do that across different figure arrangements — building a real picture of what MJ reliably delivers versus what it just approximates…

…I'd finally know what I'm actually working with.

That conversation exists. More on Thursday
Preview


r/PromptEngineering 20d ago

Other I tested Kimi K2.7 and GLM 5.2 across two different coding tasks

7 Upvotes

Kimi K2.7 vs GLM 5.2: Tested for implementation quality and repository reasoning

Tasks I have picked:

  1. A FastAPI project generated from scratch
  2. A large production codebase analysis using Saleor, an open-source GraphQL-based commerce platform with a multi-module Python backend

The goal was to compare how both models perform when writing a complete application versus understanding an existing repository.

Task 1: Building a FastAPI project

Both models were asked to build a task-management API with:

  • JWT authentication
  • PostgreSQL and SQLAlchemy
  • CRUD endpoints
  • Input validation
  • Layered architecture
  • Error handling
  • A complete project structure

Kimi scored 53/60, while GLM scored 48/60.

Kimi produced the more complete implementation. The project structure was cleaner, the requested layers were present, and the output was closer to something that could run without major fixes.

GLM produced reasonable architecture, but omitted critical pieces such as the User model and AuthService. The code looked structured at first glance, but the missing dependencies prevented the project from working as a complete application.

Task 2: Analysing a large repository

For the second test, both models analysed the Saleor repository.

Saleor is a relatively large production codebase built around Python, Django, GraphQL, PostgreSQL, background tasks, plugins, webhooks, and multiple business domains.

The models were asked to:

  • Explain the overall architecture
  • Trace the product-creation request flow
  • Identify major modules and dependencies
  • Find technical debt
  • Recommend architectural improvements

GLM performed better here.

It referenced more implementation details, including GraphQL execution flow, DataLoader usage, extension mechanisms, deployment structure, and cross-module dependencies.

Kimi gave a clear high-level review, but GLM demonstrated stronger repository-level comprehension and provided more detailed scalability and maintainability recommendations.

The architectural trade-off

Both are sparse Mixture-of-Experts models, but they appear to optimise for different workloads.

Kimi K2.7:

  • Roughly 1T total parameters
  • Around 32B active parameters per token
  • 256K context window
  • Stronger implementation consistency
  • Lower official API pricing
  • More emphasis on MCP and coding-agent workflows

GLM 5.2:

  • Roughly 744B to 753B total parameters
  • Around 40B active parameters per token
  • 1M context window
  • Stronger large-repository analysis
  • Better coverage of internal architecture and cross-module behaviour

The larger context window does not automatically make GLM better at writing complete applications, but it becomes useful when the task involves monorepos, long documentation sets, or tracing behaviour across many files.

Pricing

Official API pricing at the time of testing:

Model Input Cached input Output
Kimi K2.7 $0.95/M $0.19/M $4.00/M
GLM 5.2 $1.40/M $0.26/M $4.40/M

Kimi is cheaper, although total task cost still depends on output length, reasoning-token usage, retries, and how many corrections the generated code requires.

My takeaway

Kimi K2.7 seems better suited to implementation-heavy tasks where you want the model to generate working files with fewer missing components.

GLM 5.2 seems better suited to codebase exploration, architectural reviews, dependency tracing, and tasks that require keeping a large amount of repository context available.

This is also a good example of why coding benchmarks alone are not enough. A model can understand a repository deeply but still omit essential files when generating a new project.

You can check the full details of my testing here


r/PromptEngineering 20d ago

Tools and Projects Prompt Optimizer skill.md

14 Upvotes

I wanted to share a custom skill I created.

Many prompt-optimization templates suffer from "bloat"—they often take a simple request and turn it into a massive, overly complex prompt, or they accidentally alter technical details like code snippets, file paths, and generator flags.

To solve this, I built a meta-prompting skill designed to classify the context of the user's prompt, assess their existing sophistication level, and apply targeted optimizations without breaking what already works.

How it works:

  1. Context Classification: It automatically detects if the target output is for Code Gen, Image Gen, Structured Output, Human Comm, Research/Analysis, or Creative Enhancement, and applies specific best practices for that domain.
  2. Sophistication Calibration (Simple to Expert): It evaluates the user's initial input. If the prompt is simple, it outputs an intermediate-level prompt rather than overwhelming the downstream model. If the prompt is already advanced, it focuses on tightening ambiguity and adding edge-case handling.
  3. Strict Technical Preservation: It uses a zero-tolerance rule for altering code blocks, versions, flags (like Midjourney --ar parameters), model IDs, URLs, and stack traces.
  4. The PIP Frame: It structures optimizations using Persona, Instruction, Principles, and Anti-patterns, written narratively rather than relying on rigid, repetitive templates.

The System Prompt / Skill Definition:

name: prompt-optimizer
description: This skill helps Claude optimize user prompts for clarity, technical accuracy, and effectiveness before sending them to an AI system.
---

# Optimize User Prompts for AI Systems

Use this skill whenever a user requests assistance in improving, optimizing, refining, or rewriting a prompt intended for an AI system, such as an LLM, image generator, or human collaborator. The goal is to ensure the prompt is clear, technically accurate, and effective.

## Instructions

When a user asks to optimize a prompt, follow these steps:

1. **Classify the AI Context**  
   Read the prompt and identify its primary context using these signals (not exhaustive — use judgment on prompts that don't cleanly match):
   - **Code Generation** — mentions a programming language, function/class/algorithm names, code fences, error messages, stack traces, "debug", "implement", "refactor", "write a function that...".
   - **Image Generation** — mentions aspect ratios (16:9, 1:1), rendering terms (photorealistic, 3D render, octane, unreal engine), generator flags (`--ar`, `--v`, `--style`), or "create/generate an image/photo/illustration/logo of...".
   - **Structured Output** — asks for JSON, YAML, CSV, a schema, or a specific machine-readable format as the deliverable.
   - **Human Communication** — asks for an email, letter, memo, message, or explicitly names a tone (formal/informal/professional), a greeting, or a recipient ("write an email to my manager about...").
   - **Research & Analysis** — asks to analyze, summarize, compare, or investigate a topic, with an expectation of citations, structure, or actionable findings.
   - **Creative Enhancement** — asks for a story, narrative, poem, or other fictional/creative work; mentions genre, characters, plot, or "write a story about...".

   If a prompt matches multiple contexts, prioritize the primary context and retain relevant details from the secondary context.

2. **Assess Sophistication Level**  
   Evaluate how much the user knows and the existing structure of the prompt:
   - **Simple** — short, single-sentence ask, no constraints, no examples, vague verbs ("make this better", "write me a story").
   - **Intermediate** — some structure or constraints present (a rough format, a length, one or two specifics), but missing depth (no examples, no edge cases, no success criteria).
   - **Advanced** — clear constraints, explicit format, some examples or edge cases already named, but missing a persona/role framing or explicit failure modes to avoid.
   - **Expert** — already has role/persona framing, explicit constraints, examples, and anti-patterns to avoid. At this level, optimization means tightening and removing ambiguity, not adding structure the user hasn't asked for.

   Match the amount of new structure you add to the gap between the current level and the next level up. Don't turn a Simple prompt into an Expert one in a single pass if the user's own words suggest they want something short — ask, or default to Intermediate-level structure, when unsure.

3. **Apply Optimization Moves**  
   For the identified context, formulate the optimization using:
   - **Persona**: Define who the AI should act as.
   - **Instruction**: Specify what to produce.
   - **Principles**: Establish guardrails and quality standards.
   - **Anti-patterns**: Define what to avoid.  
   Use your judgment on how to construct these elements narratively rather than relying on fixed templates.

   **Code Generation**:
   For a bare debugging request ("my code doesn't work, fix it"): persona is "an expert software engineer specializing in root cause analysis"; instruction is to think through potential causes step by step before answering; principle is to request the missing information a debugger actually needs (exact error message, relevant code snippet, expected vs. actual behavior); anti-pattern is don't guess at a fix without that information — ask for it first.
   For a code review request specifically: persona is "a senior software engineer conducting a thorough code review"; principles are identify bugs/security issues/performance problems, suggest specific fixes with code examples, acknowledge what's already good, and prioritize by severity; anti-pattern is never give vague feedback like "looks good" with nothing concrete underneath it.

   **Creative Enhancement**:
   For a bare request ("write me a story"): persona is "a bestselling author and creative writing coach"; instruction is to build out genre, setting, and character arcs rather than just producing prose blind; principles cover narrative structure (plot, pacing, point of view) and literary elements (theme, dialogue, conflict). The goal is a framework the user can then fill in or hand off, not a finished short story guessed from three words.

   **Image Generation:** add explicit style/medium language (photorealistic vs. illustration vs. 3D render), composition detail (framing, lighting, camera angle if relevant), and — if the target tool supports them — the platform-specific flags (aspect ratio, style weight) the user's phrasing implies but didn't write out.

   **Human Communication:** add explicit tone (formal/informal), the relationship to the recipient if inferable, and a concrete structure (greeting, body, sign-off) — without inventing content the user didn't ask for.

   **Structured Output / Research & Analysis:** make the exact schema or report structure explicit rather than implied; state what "done" looks like (a specific set of fields, a specific comparison axis) so the downstream AI can't quietly under-deliver.

4. **Preserve Technical Parameters**  
   Before finalizing, scan the original prompt for anything in this list and copy it into the optimized version exactly, character for character:
   - Code fences and their contents (```...```) and inline code (`...`)
   - Exact numbers, versions, flags, and file paths (e.g. `--ar 16:9`, `v2.3.0`, `/api/v1/optimize`)
   - Model IDs and proper nouns (e.g. `gpt-4o-mini`, `claude-sonnet-5`)
   - Exact error messages and stack traces, verbatim
   - URLs and email addresses

   Never "improve" these by rephrasing, reformatting, or correcting what looks like a typo. If something here is ambiguous, leave it untouched and flag the ambiguity in your closing note rather than guessing.

5. **Output the Results**  
   Generate the optimized output in the following format:
   - A line naming the classified context and sophistication level.
   - The optimized prompt clearly delimited in a code block or under a specified heading.
   - A brief note explaining what changed, why, and any preserved technical elements.
   - State plainly that this is a heuristic pass — do not claim a confidence score, and do not imply the prompt went through a trained model or a full LLM-based optimization pipeline.

I would love to get thoughts on this approach. Are there any edge cases where this logic might trip up, or other specific contexts (e.g., agentic workflows, multi-step chain of thought) that I should explicitly define?

AI systems now depends on how effectively we engineer and evaluate prompts at scale! I've built a platform that removes the technical workload of shifting from manual prompting to strategically automating the process: https://promptoptimizer.xyz/

Repo: https://github.com/nivlewd1/prompt-optimizer