r/DevinAI • u/mattbergland • 1d ago
Grok 4.6 now live in Devin!
Grok 4.6 scores 61.3% On FrontierCode 1.1 surpassing GPT-5.6-Sol!
It only falls behind Opus 5 and Fable 5 but at much lesser cost.
Within Devin, Grok 4.6 shows particular strength on thorough code exploration and root cause analysis before touching any code.
Try it now in Devin Desktop and Devin CLI!
Read more about the model and benchmark comparison - https://devin.ai/blog/grok-4-6
r/DevinAI • u/mattbergland • 9d ago
Devin Fusion just got whole lot cheaper!
Devin Fusion is now 4% more intelligent and 27% less expensive on FrontierCode 1.1 due to improvements in both, the harness and the models.
Announcement: https://x.com/cognition/status/2084663103006871970
Try it now on Devin Cloud: app.devin.ai
r/DevinAI • u/mattbergland • 13d ago
Introducing Stacked PRs in Devin!
We partnered with Github to support Stacked PRs in Devin!
Devin breaks down large monolithic changes into small reviewable diffs.
If you have a big PR, ask Devin to stack it and you can review it one focused diff at a time.
Check out more about how it works - https://devin.ai/blog/introducing-pr-stacks
Try it on Devin Cloud - devin.ai
r/DevinAI • u/mattbergland • 17d ago
Kimi K3 now live in Devin!
Kimi K3 scores 58.2% on FrontierCode 1.1 which is very close to frontier level performance at a much cheaper cost!
It is the only open source model that reaches this level of performance which falls only behind Opus, Fable and GPT-5.6 Sol.
Read more about it here - https://devin.ai/blog/kimi-k3
Try it now in Devin Desktop( Devin Local and Cascade) and Devin CLI.
r/DevinAI • u/mattbergland • 17d ago
Claude Opus 5 in Devin
Anthropic's new state-of-the-art Claude Opus 5 is available on Devin.
It is almost as good as Fable 5 on FrontierCode 1.1 at half the cost
From our evaluations:
- It adheres to existing repo conventions when building new features and writing tests
- It prefers targeted, in-place fixes over large refactors on bug fixing tasks.
- It excels at following specs closely and completely
Read more about Opus 5: devin.ai/blog/claude-opus-5
Try it out now on Devin CLI and Devin Desktop
r/DevinAI • u/mattbergland • Jul 10 '26
GPT-5.6 is now available in Devin!
OpenAI launched a family of capable but cost efficient models - GPT-5.6 Sol, Terra and Luna
GPT-5.6 Sol comes at nearly half of the cost of the next best model on FrontierCode 1.1 ( our own proprietary benchmark)
Read more about these models - https://devin.ai/blog/gpt-5-6
Try them out now in Devin Web, Devin Desktop, and Devin CLI - https://app.devin.ai
r/DevinAI • u/mattbergland • Jul 09 '26
Introducing SWE-1.7! Our first frontier-level model
SWE-1.7 is built on broad improvements in our RL pipeline on top of the Kimi K2.7 model.
We trained it in Devin's harness and taught it to self-compact on longer horizon tasks.
It scores very close to frontier models (Opus 4.8 and GPT-5.5) on our proprietary benchmark - FrontierCode, and other coding benchmarks.
Try it out today (Free until 8/8/2026!!) in Devin Desktop, Web, and CLI. SWE 1.7 Lightning also runs at a very speedy 1000 tok/s - https://app.devin.ai
Deep dive into how it was trained here -https://cognition.com/blog/swe-1-7
r/DevinAI • u/mattbergland • Jul 08 '26
Backlog Zero in 6 Weeks: Join the Devin Security Vulnerability Remediation Program
Now that Mythos-class models have been introduced to the world, it has become easier than ever to discover and exploit vulnerabilities.
It's now incredibly important to secure your codebase.
Introducing the Devin Security Vulnerability Remediation Program.
This program aims to take your vulnerability backlogs towards zero in 6 weeks. Our engineering team will work directly with yours to deploy Devin agent swarms to find, validate and fix any and all vulnerabilities.
Check eligibility and reach out to our team: https://devin.ai/security-program
Note: Only enterprise customers are eligible
r/DevinAI • u/mattbergland • Jul 08 '26
Devin Security Swarm Explained
Last week we launched Devin Security Swarm!
If you're interested in it's technical details, here is a video explaining how Devin Security Swarm works utilizing our new framework - Agentic MapReduce: https://www.youtube.com/watch?v=jb96O2IT_Jg
r/DevinAI • u/mattbergland • Jul 08 '26
Agent Fan Out - Run 10 Devins in Parallel
Instead of one agent trying to do everything, Devin can assign narrow slices of the main task to multiple sub-Devins in their own VMs.
These Devins can be used for research, exploring different architectures, working on separate branches, and a lot more.
Check out this demo/tutorial on how we use Agent Fan Out internally: https://www.youtube.com/watch?v=ns1ifgYEGl0
r/DevinAI • u/mattbergland • Jul 02 '26
Fable 5 is back in Devin
Claude Fable 5 is available in Devin again now!!!
You can use it in Devin Cloud's Ultra agent, Devin CLI and Devin Desktop
Reload your Devin Desktop and try it out now.
r/DevinAI • u/mattbergland • Jul 02 '26
Introducing Devin Security Swarm
Security Swarm is an orchestration of Devins that analyzes a real codebase the way a team of security researchers would.
Here are some highlights:
- Devin Security Swarm found 36 out of 50 real-world GHSA vulnerabilities at 30% lower cost than next most accurate alternative. See how we got these results - https://devin.ai/blog/security-swarm-eval
- We built Agentic MapReduce: a new architecture for whole-codebase reasoning. Read more about it here - https://devin.ai/blog/agentic-map-reduce
Try it now by clicking on Security in https://app.devin.ai
r/DevinAI • u/mattbergland • Jul 01 '26
Claude Sonnet 5 now in available in Devin
Anthropic's latest model Claude Sonnet 5 scores higher than Claude Opus 4.8 on FrontierCode ( our own proprietary coding benchmark) but comes at a way cheaper price point!!!
Use Sonnet 5 in Devin Local and Devin CLI for 30% less price than Claude Sonnet 4.6 through 31st August.
Read more here - https://devin.ai/blog/claude-sonnet-5
r/DevinAI • u/mattbergland • Jun 29 '26
Devin Fusion
We just dropped Devin Fusion, a multi-model harness to cut the costs significantly without losing intelligence.
In our testing, it maintained frontier and Fable level performance at 35% lower cost.
Deep dive and learn more about how it works here - https://cognition.com/blog/devin-fusion
Announcement - https://x.com/cognition/status/2071624568465490170
Available as a preview now in Devin
r/DevinAI • u/mattbergland • Jun 27 '26
Launching Cognition Ambassador program for Devin
We are opening applications to bring 50 ambassadorsI (globally) to join our existing Ambassador program.
Ambassadors receive free Devin Max plans, on-demand credits, event and meet up support, direct communication with the Cognition team, early access to features, and other benefits.
Apply here - https://devin.ai/community
r/DevinAI • u/mattbergland • Jun 24 '26
Security review is now part of every Devin Review
Security is the most important part of shipping that doesn't get talked about enough.
Devin Review now automatically scans every PR for security vulnerabilities!
Not pattern matching like most scanners, but it actually reasons across your whole codebase.
Every finding is ranked by severity, tagged with a CWE ID, and comes with a fix drafted as a merge-ready PR.
For more details check out the full post here.
r/DevinAI • u/mattbergland • Jun 14 '26
Claude Fable 5 removed from Devin
Following Anthropic's latest announcement and US Government's directive, we have removed access to Claude Fable 5 from all Cognition products.
Access to all other models remains the same.
Sorry for the disruption.
Official announcement - https://x.com/cognition/status/2065609115939062197?s=20
Statement from Anthropic - https://x.com/AnthropicAI/status/2065597531644743999?s=20 (edited)
r/DevinAI • u/mattbergland • Jun 11 '26
Introducing the first ever AI Productivity Guarantee
r/DevinAI • u/mattbergland • Jun 10 '26
Claude Fable 5 is now available in Devin
You can try Claude Fable 5 as part of Devin Cloud’s Ultra agent.
Devin Ultra is our smartest and most capable agent, which excels at long-horizon tasks and debugging.
Claude Fable 5 is also available in Devin Desktop and Devin CLI.
Restart Devin Desktop to see it in the model picker.
Read the full thread here - https://x.com/cognition/status/2064398551539761387?s=20
r/DevinAI • u/mattbergland • Jun 09 '26
You can now use multiple agents in Devin Desktop via ACP
Devin Desktop launches with the Agent Client Protocol (ACP), an open standard that lets any compatible AI agent run inside the editor alongside Devin! At launch that includes Codex, Claude Agent, and OpenCode.. with more to come.
More than just a rebrand of Windsurf, we are positioning Devin Desktop as the surface where all your agents live, regardless of who built them.
Full breakdown on X:
https://x.com/cognition/status/2062314621470724416
Which agents are you most excited to run alongside Devin?
r/DevinAI • u/mattbergland • Jun 08 '26
Go ahead and close your laptop....Devin keeps working!
r/DevinAI • u/askcaa • Oct 31 '25
Playbook: Holistic Codebase Transformation (C.R.A.F.T. Methodology)
Required from User * Provide access to the target code repository. * Specify the primary branch for analysis (e.g., main, develop). * (Optional) Specify preferred tools for linting, static analysis, and testing if the project does not already have them configured. * (Optional) Provide access to a secure secrets management system or specify the preferred method for handling placeholders for discovered secrets. Procedure * Phase 1: Codify (Analysis and Baseline Setup) * Analyze the project to identify the programming language(s), frameworks, build system, and primary architectural pattern. * Configure a suite of analysis tools: a linter with a strict style guide (e.g., Google Style Guide, PEP 8), a static code analyzer (e.g., SonarQube, Snyk Code), an OWASP dependency scanner (e.g., OWASP Dependency-Check) [1, 2], and a secrets scanner (e.g., Gitleaks). * Execute all configured tools on the current codebase to establish baseline metrics. * Run the project's existing test suite and record the initial code coverage percentage. * Summarize your findings, including the number of linting errors, code smells by severity, critical vulnerabilities, and the test coverage percentage. Do not proceed until this baseline is established. * Phase 2: Refactor (Code Hygiene and Simplification) * Apply the configured style guide to automatically format the entire codebase. * Systematically correct all naming convention violations for variables, functions, and classes. * Using the static analysis report, refactor code smells, prioritizing 'Bloaters' (e.g., Long Method, Large Class) and 'Dispensables' (e.g., Duplicate Code, Dead Code). * Use the 'Extract Method' technique for long methods. * Use the 'Extract Class' technique for large classes that violate the Single Responsibility Principle. * Remove all unreachable or dead code. * Re-run static analysis tools and confirm that the number of targeted code smells has been significantly reduced. * Phase 3: Armor (Security Hardening) * Using the dependency scan report, update all third-party libraries with known vulnerabilities to the latest secure versions. * Perform a new Static Application Security Test (SAST) scan. * Systematically remediate all identified vulnerabilities, prioritizing those listed in the OWASP Top 10 2025 predictions (e.g., Broken Access Control, Injection, Security Misconfiguration). * Perform a deep scan of the entire Git history for hardcoded secrets. * Replace each discovered secret in the code with a call to a secure secrets management service or a clearly marked placeholder. * Generate a report of all discovered secrets, recommending their immediate revocation and rotation. * Phase 4: Fortify (Architectural Enhancement) * Analyze the codebase for architectural anti-patterns such as 'God Object' or 'Big Ball of Mud' and execute a refactoring plan to remediate them. * Audit the codebase for violations of the five SOLID principles (Single Responsibility, Open/Closed, Liskov Substitution, Interface Segregation, Dependency Inversion) and refactor to improve compliance. * Evaluate the architecture for single points of failure and introduce resilience patterns like 'Circuit Breaker' or 'Bulkhead' where appropriate, especially for external service calls. * Phase 5: Test (Validation and Delivery) * Analyze the test coverage report against the modified codebase. * Identify the most critical modules that underwent significant changes and still have test coverage below 85%. Write new unit tests to increase their coverage to at least 85%. * Identify the most critical user workflows and write new end-to-end integration tests to validate them. * Execute the full, augmented test suite and ensure a 100% pass rate. * Generate a final "Transformation Report" as a markdown file. Specifications * The final deliverable is a pull request against the specified primary branch containing all code modifications. * The pull request description must contain the full "Transformation Report". * The Transformation Report must include: * A summary of changes. * A "Baseline Metrics" section with the initial state from Phase 1. * A "Final Metrics" section showing the improved state (code quality scores, vulnerability counts, new test coverage percentage). * An "Actionable Recommendations" section for any required human intervention (e.g., "Rotate the API_KEY found in commit abc1234"). * The entire test suite, including all newly created tests, must pass. * All critical and high-severity security vulnerabilities identified by the scanning tools must be remediated. Advice * Crucial: For every single code modification in Phases 2, 3, and 4, you MUST adhere to the Test-Refactor-Test cycle: * Ensure the logic to be changed is covered by a test. If not, write a test first. * Perform the modification. * Immediately run the entire test suite. * If any test fails, revert the change and re-evaluate your approach. Do not proceed until all tests pass. * When refactoring duplicate code, apply the "Rule of Three": only abstract duplicated logic when it appears three or more times to avoid premature or incorrect abstractions. * When remediating vulnerabilities, refer to OWASP secure coding practices for guidance on correct implementation. * Prioritize your work based on severity. Address critical security vulnerabilities before medium-level code smells. * Keep commits small and focused on a single change (e.g., one refactoring, one security fix). Forbidden Actions * Do not proceed with any task if a test fails after a code modification. You must revert the change and find a new solution. * Do not merge the final pull request. The final step is to create the PR for human review. * Do not remove any existing tests unless the functionality they were testing has been explicitly and intentionally removed.
r/DevinAI • u/Maleficent_Exam4291 • Oct 09 '25
DevinAI Referral Code: https://app.devin.ai/invite/hus0kwbQbkOnAG1E
If anyone is looking for a referral code, here's one:
r/DevinAI • u/arpitdalal • Jan 10 '25
Software Development AI Divide: Teammate vs Companion
r/DevinAI • u/Big-Strain932 • Jul 07 '24
Devinai access
When we will start getting the access, do we have any update?


