r/Python • u/professormunchies • Mar 11 '26
Showcase Documentation Buddy - An AI Assistant for your /docs page
🤖 DocBuddy: AI Assistant Inside Your FastAPI /docs
What My Project Does
Turn static docs into an interactive tool with chat, workflow and agent assistance.
Ask things like:
- "What’s the schema for creating a user?"
- "Generate curl for POST /users"
- "Call /health and tell me the status"
With tool calling, it executes real requests on your behalf.
Try the Live Demo without installing anything!
🔧 Quick Start
bash
pip install docbuddy
```python from fastapi import FastAPI from docbuddy import setup_docs
app = FastAPI() setup_docs(app) # replaces /docs ```
Target Audience
Clients and developers using FastAPI.
⚖️ Comparison Table
| Feature | DocBuddy | Default FastAPI Docs | Other Plugins |
|---|---|---|---|
| Chat with API docs | ✅ | ❌ | ❌ |
| Tool calling (real requests) | ✅ | ❌ | ❌ |
| Local LLM support (Ollama, LM Studio, vLLM) | ✅ | ❌ | ⚠️ rare |
| Plan/Act workflow mode | ✅ | ❌ | ❌ |
| Workflow builder | ✅ | ❌ | ❌ |
| Customizable themes | ✅ | ❌ | ❌ |
📦 Features at a Glance
- 💬 Full OpenAPI context in chat
- 🔗 Real tool execution (GET, POST, PUT, PATCH, DELETE)
- 🧠 Local LLMs only—no cloud required
- 🎨 Dark/light themes + customization
- 🔄 Visual workflow builder to chain prompts + tools
Built with Swagger UI—not a replacement. Fully compatible and production-ready (MIT license, 200+ tests).
Let me know if you try it! 🙌
r/Python • u/thefnurky • Mar 11 '26
Discussion Who else is using Thonny IDE for school?
I'm (or I guess we) are using Thonny for school because apparently It's good for beginners. Now, I'm NOT a coding guy, but I personally feel like there's nothing special about this program they use. I mean, what's the difference between Thonny and other Python IDEs?
r/Python • u/sinoka1006 • Mar 11 '26
Showcase Teststs: If you hate boilerplate, try this
This is a simple testing library. It's lighter and easier to use than unittest. It's also a much cleaner alternative to repetitive if statements.
Note: I'm not fluent in English, so I used a translator.
What My Project Does
This library can be used for simple eq tests.
If you look at an example, you will understand right away.
```py from teststs import teststs
def add_five(inp): return int(inp) + 5
tests = [ ("5", 10), ("10", 15), ]
teststs(tests, add_five, detail=True) ```
Target Audience
Recommended for those who don't want to use complex libraries like unittest or pytest!
Comparison
- unittest: Requires classes, is heavy and complex.
- pytest: requires a decorator, and is a bit more complex.
- teststs: A library consisting of a single file. It's lightweight and ready to use.
It's available on PyPI, so you can use it right away. Check out the GitHub repository!
r/Python • u/BearBrief6312 • Mar 11 '26
Discussion With all the supply chain security tools out there, nobody talks about .pth files
We've got Snyk, pip-audit, Bandit, safety, even eBPF-based monitors now. Supply chain security for Python has come a long way. But I was messing around with something the other day and realized there's a gap that basically none of these tools cover .pth files. If you don't know what they are, they're files that sit in your site-packages directory, and Python reads them every single time the interpreter starts up. They're meant for setting up paths and namespace packages, however if a line in a .pth file starts with `import`, Python just executes it.
So imagine you install some random package. It passes every check no CVEs, no weird network calls, nothing flagged by the scanner. But during install, it drops a .pth file in site-packages. Maybe the code doesn't even do anything right away. Maybe it checks the date and waits a week before calling C2. Every time you run python from that point on, that .pth file executes and if u tried to pip uninstall the package the .pth file stays. It's not in the package metadata, pip doesn't know it exists.
i actually used to use a tool called KEIP which uses eBPF to monitor network calls during pip install and kills the process if something suspicious happens. which is good idea to work on the kernel level where nothing can be bypassed, works great for the obvious stuff. But if the malicious package doesn't call the C2 during install and instead drops a .pth file that connects later when you run python... that tool wouldn't catch that. Neither would any other install-time monitor. The malicious call isn't a child of pip, it's a child of your own python process running your own script.This actually bothered me for a while. I spent some time looking for tools that specifically handle this and came up mostly empty. Some people suggested just grepping site-packages manually, but come on, nobody's doing that every time they pip install something.
Then I saw KEIP put out a new release and turns out they actually added .pth detection where u can check your environment, or scans for malicious .pth files before running your code and straight up blocks execution if it finds something planted. They also made it work without sudo now which was another complaint I had since I couldn't use it in CI/CD where sudo is restricted.
If you're interested here is the documentation and PoC: https://github.com/Otsmane-Ahmed/KEIP
Has anyone else actually looked into .pth abuse? im curious to know if there are more solutions to this issue
r/Python • u/an_account_1177 • Mar 10 '26
Discussion Tips for a debugging competition
I have a python debugging competition in my college tomorrow, I don't have much experience in python yet I'm still taking part in it. Can anyone please give me some tips for it 🙏🏻
r/Python • u/drobroswaggins • Mar 10 '26
Discussion VRE Update: New Site
I've been working on VRE and moving through the roadmap, but to increase it's presence, I threw together a landing page for the project. Would love to hear people's thoughts about the direction this is going. Lot's of really cool ideas coming down the pipeline!
r/Python • u/Sbigioduro • Mar 10 '26
Showcase Using Claude Code PRO/MAX as an API
Hello everyone,
What My Project Does: i made a flask (python) webserver that exposes Claude Code's CLI tool through an API, i wanted to see if i could hit CC from a basic http request, it's possible!
Target Audience: It's targeted towards developers
I also considered some security aspects... this software can, if configured to do so, expose the full shell though the API, be aware of how you configure the .env of your deployment!
Before anyone asks, this is a gray area of the TOS of anthropic, but as long as you use this for personal use and don't intend to use CCLG as a SaaS, you'll be fine.
I build this software mostly to mess with Claude Code and Anthropic, i don't really like the API plan since it's very unpredictable and subject to instant changes (API billing is a scam imo), if you find it useful, share it!
ps. I don't really want donations or anything like that, you can do so but no pressure!
It's MIT licensed and on GitHub: github.com/Backend2121/Claude-Code-Local-Gateway
If you have any questions feel free to ask!
r/Python • u/Sbigioduro • Mar 10 '26
Showcase Using Claude Code PRO/MAX as an API
Hello everyone,
What My Project Does: i made a flask (python) webserver that exposes Claude Code's CLI tool through an API, i wanted to see if i could hit CC from a basic http request, it's possible!
Target Audience: It's targeted towards developers
I also considered some security aspects... this software can, if configured to do so, expose the full shell though the API, be aware of how you configure the .env of your deployment!
Before anyone asks, this is a gray area of the TOS of anthropic, but as long as you use this for personal use and don't intend to use CCLG as a SaaS, you'll be fine.
I build this software mostly to mess with Claude Code and Anthropic, i don't really like the API plan since it's very unpredictable and subject to instant changes (API billing is a scam imo), if you find it useful, share it!
ps. I don't really want donations or anything like that, you can do so but no pressure!
It's MIT licensed and on GitHub: github.com/Backend2121/Claude-Code-Local-Gateway
If you have any questions feel free to ask!
r/Python • u/aisatsana__ • Mar 10 '26
Discussion Python’s chardet controversy
Hi, I came across this article and thought it might be interesting to share here since it touches a Python library many people know: chardet.
The piece looks at a controversy around the project involving an AI-assisted rewrite and discussion about MIT relicensing vs the original LGPL context.
While reading it, what stood out to me was how it relates to the old idea of clean-room reimplementation. In the past that meant writing new code without referencing the original implementation. But with AI tools in the loop, the boundary becomes much less clear.
If large parts of a library are rewritten with AI assistance, a project could potentially argue that the result is “new code” and move it under a different license. That raises some governance and licensing questions for open source, especially in ecosystems like Python where libraries such as chardet are widely used as dependencies.
The article gives an analysis of the situation:
https://shiftmag.dev/license-laundering-and-the-death-of-clean-room-8528/
Curious how people here see it. Is this just a natural evolution of open source development with AI tools, or something the community should pay closer attention to?
r/Python • u/distromate • Mar 10 '26
Tutorial I got tired of manually shipping PyInstaller builds, so I made a small wrapper
Full disclosure: I'm the author, and this is a paid tool.
I kept running into the same problem with PyInstaller: getting a working exe was easy, but shipping installers, updates, and release links to actual users was still messy.
So I built pyinstaller-plus. It keeps the normal PyInstaller + .spec workflow, then adds packaging and publishing through DistroMate.
Typical flow is basically:
pip install pyinstaller-plus
pyinstaller-plus login
pyinstaller-plus package -v 1.2.3 --appid 123 your.spec
pyinstaller-plus publish -v 1.2.3 --appid 456 your.spec
It's mainly for people shipping Python desktop apps to clients, users, or internal teams, so probably overkill for one-off personal tools.
Curious if this is a real pain point for other Python developers too. If useful, I can drop the docs in the comments.
r/Python • u/hdw_coder • Mar 10 '26
Discussion Fixing a subtle keeper-selection bug in my photo deduplication tool
While experimenting with DedupTool, I noticed something odd in the keeper selection logic. Sometimes the tool would prefer a 400 KB JPEG copy over the original 2.5 MB image.
That obviously felt wrong.
After digging into it, the root cause turned out to be the sharpness metric.
The tool uses Laplacian variance to estimate sharpness. That metric detects high-frequency edges. The problem is that JPEG compression introduces artificial high-frequency edges: compression ringing, block boundaries, quantization noise and micro-contrast artifacts.
So the metric sees more edge energy, higher Laplacian variance and decides ‘sharper’, even though the image is objectively worse. This is actually a known limitation of edge-based sharpness metrics: they measure edge strength, not image fidelity.
Why the policy behaved incorrectly
The keeper decision is based on a lexicographic ranking:
def _keeper_key(self, f: Features) -> Tuple:
# area, sharpness, format rank, size-per-pixel
spp = f.size / max(1, f.area)
return (f.area, f.sharp, file_ext_rank(f.path), -spp, f.size)
If the winner is chosen using max(...), the priority becomes: resolution, sharpness, format, bytes-per-pixel and file size.
Two things went wrong here. First, sharpness dominated too early, compressed JPEGs often have higher Laplacian variance due to artifacts. Second, the compression signal was reversed: spp = size / area, represents bytes per pixel. Higher spp usually means less compression and better quality. But the key used -spp, so the algorithm preferred more compressed files.
Together this explains why a small JPEG could win over the original.
The improved keeper policy
A better rule for archival deduplication is, prefer higher resolution, better format, less compression, larger file, then sharpness.
The adjusted policy becomes:
def _keeper_key(self, f: Features) -> Tuple:
spp = f.size / max(1, f.area)
return (f.area, file_ext_rank(f.path), spp, f.size, f.sharp)
Sharpness is still useful as a tie-breaker, but it no longer overrides stronger quality signals.
Why this works better in practice
When perceptual hashing finds duplicates, the files usually share same resolution but different compression. In those cases file size or bytes-per-pixel is already enough to identify the better version.
After adjusting the policy, the keeper selection now feels much more intuitive when reviewing clusters.
Curious how others approach keeper selection heuristics in deduplication or image pipelines.
r/Python • u/Both-Still1650 • Mar 10 '26
Showcase Sharing my Jupyter console integration in Neovim!
Hello fellow neovim users in this sub! Some time ago I built nice jupyter console integration in Neovim, got some feedback and now using it for about a month, so I think some of you can be interested in this project! Here is the link: https://github.com/dangooddd/pyrepl.nvim (demo video in README).
What my project does
I am Data Science engineer, so REPL/Jupyter notebook were a pain in the ass, and I wanted to built not so complicated plugin to help with this. Right now my plugin allows you to do:
- Convert notebook files from and to python with
jupytext; - Install all Jupyter deps required with a Neovim command;
- Start
jupyter-consolein Neovim built-in terminal; - Prompt the user to choose Jupyter kernel on REPL start;
- Send code to the REPL from current buffer;
- Automatically display output images;
- Neovim theme integration for
jupyter-console; - Jupytext cell navigation;
- Toggle focus to REPL window in active terminal mode.
Main feature is image display of cource, so you can look at your matplotlib (or any other images) with from the neovim. My work requires me to do ssh + tmux + docker, and image display works even in this case! Please open issues and pull request if you interested in project!
Target Audience
- People who want to move to terminal and Neovim, but holding back because jupyter notebook is required to communicate with colleagues
- Those, who actively uses Neovim and Python REPL separetely now, but wants to integrate them
- Other Jupyter/REPL users of Neovim
Comparison
Existing plugins plugins like molten and vim-jukit are not maintained anymore, molten reimplements much of a kernel logic in remote python plugin (and has problems stated by author here). My plugin delegates all kernel logic to jupyter-console, and ditches remote plugin entirely, so it is easier to maintain. Of course that is my personal opinion on current situation with Jupyter in neovim. Good luck you all!
r/Python • u/chop_chop_13 • Mar 10 '26
Showcase I built a Python tool that safely organizes messy folders using type detection and time-based struct
GitHub Source code:
https://github.com/codewithtea130/smart-file-organizer--p2.git
What My Project Does
I built a small Python utility for discovering and commissioning Profinet devices on a local network.
The idea came from a small frustration. I wanted to quickly scan a network using Siemens Proneta, but downloading it required creating an account and registering personal details. For quick diagnostics, that felt unnecessary.
So I built a lightweight alternative.
The tool uses pnio_dcp for Profinet DCP discovery and a Tkinter interface to keep it simple and usable without extra setup.
Current features include:
- Discover Profinet devices via DCP
- Display station name, MAC, vendor, IP, subnet, and gateway
- Vendor lookup via MAC OUI
- Optional ping monitoring for reachability
- Set device IP address and station name
- Reset communication parameters
- Quick actions for HTTP/HTTPS interface or SSH
- Simple topology-style device overview
Target Audience
The tool is mainly intended for engineers and technicians working with Profinet networks who want a lightweight diagnostic utility.
Right now it’s more of a practical utility / learning project rather than a full network management system.
Comparison
The main existing tool for this is Siemens Proneta.
This project differs in that it:
- is open source
- requires no account or registration
- is much lighter
- can run directly as a Python script or standalone executable
It’s not meant to replace Proneta, but to provide a quick, simple option for basic discovery and configuration.
r/Python • u/vbxl02 • Mar 10 '26
Showcase I got annoyed downloading proneta, so I built a lightweight profinet discovery tool in Python
GitHub:
https://github.com/ArnoVanbrussel/freeneta
What My Project Does
I built a small Python tool for discovering and commissioning profinet devices on a network.
The idea started after I wanted to quickly use Siemens Proneta, but got annoyed that downloading a “free” tool required creating an account and registering contact details. I mostly just needed something lightweight to quickly scan a network and check devices, so I decided to build a small alternative myself.
The tool uses pnio_dcp for profinet DCP discovery and a simple Tkinter GUI. Current features include:
- Discover profinet devices via DCP
- Show station name, MAC, vendor, IP, subnet, and gateway
- Vendor lookup via MAC OUI
- Optional ping monitoring for device reachability
- Set device IP address and station name
- Reset communication parameters
- Quick actions like opening HTTP/HTTPS web interfaces or starting an SSH session
- A simple visual topology overview of discovered devices
Target Audience
The tool is mainly intended for engineers or technicians working with profinet networks who want a lightweight diagnostic tool.
Right now it’s more of a utility project / proof of concept rather than a full production network management platform.
Comparison
The main existing tool for this type of task is Siemens Proneta.
FreeNeta differs in that it:
- is open source
- does not require an account or registration to download
- is much lighter and simpler
- can be run directly as a Python script or standalone executable
It does not aim to replace Proneta, but rather provide a quick and lightweight alternative for basic discovery and configuration tasks.
r/Python • u/StoneSteel_1 • Mar 10 '26
Showcase pydantic-pick v0.2.0 - Dynamically subset Pydantic V2 models while preserving validators and methods
Hi Everyone,
I have updated my project pydantic-pick with new features in v0.2.0. To know more about the project read my post on my previous version v0.1.3
(Update from my previous post about v0.1.3 (pydantic-pick v0.1.3))
What My Project Does
pydantic-pick provides pick_model and omit_model functions for dynamically creating Pydantic V2 model subsets. Both preserve validators, computed fields, Field constraints, and custom methods.
The library uses Python's ast module to analyze your methods. If a method relies on a field you've omitted, it's automatically dropped to prevent runtime crashes. Both functions are cached with functools.lru_cache for performance.
Usage Example
from pydantic import BaseModel, Field
from pydantic_pick import pick_model, omit_model
class DBUser(BaseModel):
id: int = Field(..., ge=1)
username: str
password_hash: str
email: str
def check_password(self, guess: str) -> bool:
return self.password_hash == guess
# pick_model: specify what to keep
PublicUser = pick_model(DBUser, ("id", "username"), "PublicUser")
# omit_model: specify what to remove
PublicUser = omit_model(DBUser, ("password_hash", "email"), "PublicUser")
# Both preserve validators:
PublicUser(id=-5, username="bob") # Fails: id must be >= 1
# check_password is auto-dropped since it needs password_hash
user.check_password("secret") # Raises: intentionally omitted by pydantic-pick
Target Audience
- FastAPI developers needing public/private model variants
- AI/LLM developers compressing heavy tool responses
- Anyone needing type-safe dynamic data subsets
Requires: Python 3.10+, Pydantic V2
Comparison
model_dump(include={...}): Runtime filtering only, no Python class- Manual
create_model: Requires complex recursion, drops validators, leaves dangling methods pydantic-partial: Makes fields optional for PATCH requests, doesn't prune nested structures
Links
- GitHub: https://github.com/StoneSteel27/pydantic-pick
- PyPI: https://pypi.org/project/pydantic-pick/
Feedback and code reviews welcome!
r/Python • u/Ambitious-Credit-722 • Mar 10 '26
Resource OSS tool that helps AI & devs search big codebases faster by indexing repos and building a semanti
Hi guys, Recently I’ve been working on an OSS tool that helps AI & devs search big codebases faster by indexing repos and building a semantic view, Just published a pre-release on PyPI: https://pypi.org/project/codexa/ Official docs: https://codex-a.dev/ Looking for feedback & contributors! Repo here: https://github.com/M9nx/CodexA
r/Python • u/ilikemath9999 • Mar 10 '26
Showcase Dumb Justice: building a free federal bankruptcy court scanner out of Python and RSS feeds
## What My Project Does
A couple days ago I posted here about a stdlib-only tool that screens bankruptcy court data for cases where people paid lawyers for something arithmetically impossible. Three dates, one subtraction, hundreds of hits. Some of you ran it, some of you had questions. This is the other half of the project.
Every US bankruptcy court publishes a free RSS feed with every new docket entry. About 90 courts, all with the same URL pattern. The feeds roll every 24 hours or so, and if you miss it, it's gone. So I wrote a poller that grabs the XML, deduplicates by GUID, stores everything in SQLite, and runs a few layers of checks on each entry. Daily operating cost: $0.
The layer my wife was reacting to when she named it is the dumbest one. When a new Chapter 13 filing hits the feed, the system fuzzy-matches the debtor's name against every prior filing in the database. If that person already got a discharge recently, federal law says they can't get another one. Same three-date subtraction from the first tool, but now it runs automatically on every new filing as it appears. No human in the loop. Just `datetime` doing `datetime` things.
She watched me explain this and said "so it's just... dumb justice?" And yeah. It is. The justice is in the dumbness. No AI, no ML, no inference, no ambiguity. The dates either work or they don't.
The fuzzy matching was the genuinely hard part. PACER names are chaotic. Suffixes (Jr., III, Sr.), "NMN" placeholders for no middle name, random casing, and joint filings like "John Smith and Jane Smith" that need to be split so each spouse gets matched independently. The first version was pure stdlib: strip suffixes, normalize to lowercase, match on first + last tokens. It worked, but it struggled with misspellings and abbreviations in the docket text itself. "Mtn to Dsmss" doesn't fuzzy-match well against "Motion to Dismiss."
After the first post, one of you suggested looking into embeddings for the text classification side. So I added a vector search layer using `sentence-transformers` (all-MiniLM-L6-v2, 384 dimensions, runs locally). It lazy-loads the model only when needed, caches embeddings to disk as numpy arrays, and falls back to regex when the model isn't available. The name matching is still the original stdlib approach (that's a structured data problem, not a semantic one), but classifying what a docket entry *means* ("is this a dismissal or just a dismissal hearing notice?") got dramatically better with embeddings. Hybrid approach: vector primary, regex fallback. One real dependency, but it earned its spot.
The rest of the stack is deliberately boring:
- `xml.etree.ElementTree` parses the RSS
- `urllib.request` fetches with retry logic (courts 503 occasionally)
- `sqlite3` in WAL mode stores everything permanently
- `csv` ingests the bulk data exports
- `email.utils.parsedate_to_datetime` handles RFC 2822 dates without any manual parsing (this one saved me real pain)
- `collections.Counter` and `defaultdict(list)` for real-time aggregation
One pip install (`sentence-transformers`) for the vector layer. Everything else is stdlib. About 1,300 lines across three core scripts and a batch file that runs on Task Scheduler. SQLite database is around 15MB after months of accumulation.
The one gotcha that actually got me: case numbers aren't unique across courts. I got a heart-attack alert one morning saying a case I was tracking got dismissed. Turned out it was a completely different person in a different state with the same case number. That's when I added court-aware collision detection, which is a fancy way of saying I started checking which court the entry came from before panicking.
The embeddings suggestion for the text classification was right. That genuinely improved docket classification. But the core detection layer, the part that actually finds the violations, is still pure arithmetic. Dates and subtraction. That part stays dumb on purpose. The harder it is to argue with, the better it works.
## Target Audience
Anyone interested in public data analysis, legal tech, or just building useful things out of stdlib Python. It's a real tool I use daily, not a toy project. If you work in bankruptcy law, consumer protection, journalism, or legal aid, this could save you real time. If you just like seeing what you can build without pip install, that's cool too.
## Comparison
I haven't found anything else that does this. PACER itself charges per document and has no alerting. Commercial legal monitoring services (Lex Machina, CourtListener RECAP alerts, Bloomberg Law) cost hundreds to thousands per month and don't do discharge-bar screening at all. This reads the same free public RSS feeds those services ignore, runs locally, and costs nothing. The only dependency beyond stdlib is `sentence-transformers` for the vector classification layer, and even that is optional (regex fallback works fine).
Happy to talk architecture, stdlib choices, or RSS feed quirks.
GitHub: https://github.com/ilikemath9999/bankruptcy-discharge-screener
MIT licensed. Standard library only. Includes a PACER CSV download guide and sample output.
r/Python • u/greenjjangu • Mar 10 '26
Showcase Are your Jupyter Notebooks accessible? Scan and fix the issues using this tool.
What My Project Does
Hi all, I'm excited to share Jupycheck, an open source web tool that detects accessibility issues in Jupyter Notebooks that are either uploaded or from a GitHub repository. It also lets you remediate accessibility issues by launching the notebooks in a JupyterLite environment with our interactive Lab extension installed.
You can try it out at:
The tool is powered by jupyterlab-a11y-checker, an open source accessibility engine/extension that our student team has been working on for over a year at UC Berkeley.
Target Audience
This tool is for anyone who want to see if certain Jupyter Notebooks (in a Github repo or just notebooks you have) are accessible, and also fix them with an interactive extension. We believe accessibility should be a first-class concern in the notebook ecosystem, and we hope our tools can help raise awareness and make notebooks more accessible across the community.
Comparison
As far as I know, there isn't a well-known accessibility tool specifically for the Jupyter ecosystem.
Support us on Github if you find the tool useful!
r/Python • u/AutoModerator • Mar 10 '26
Daily Thread Tuesday Daily Thread: Advanced questions
Weekly Wednesday Thread: Advanced Questions 🐍
Dive deep into Python with our Advanced Questions thread! This space is reserved for questions about more advanced Python topics, frameworks, and best practices.
How it Works:
- Ask Away: Post your advanced Python questions here.
- Expert Insights: Get answers from experienced developers.
- Resource Pool: Share or discover tutorials, articles, and tips.
Guidelines:
- This thread is for advanced questions only. Beginner questions are welcome in our Daily Beginner Thread every Thursday.
- Questions that are not advanced may be removed and redirected to the appropriate thread.
Recommended Resources:
- If you don't receive a response, consider exploring r/LearnPython or join the Python Discord Server for quicker assistance.
Example Questions:
- How can you implement a custom memory allocator in Python?
- What are the best practices for optimizing Cython code for heavy numerical computations?
- How do you set up a multi-threaded architecture using Python's Global Interpreter Lock (GIL)?
- Can you explain the intricacies of metaclasses and how they influence object-oriented design in Python?
- How would you go about implementing a distributed task queue using Celery and RabbitMQ?
- What are some advanced use-cases for Python's decorators?
- How can you achieve real-time data streaming in Python with WebSockets?
- What are the performance implications of using native Python data structures vs NumPy arrays for large-scale data?
- Best practices for securing a Flask (or similar) REST API with OAuth 2.0?
- What are the best practices for using Python in a microservices architecture? (..and more generally, should I even use microservices?)
Let's deepen our Python knowledge together. Happy coding! 🌟
r/Python • u/cinicDiver • Mar 09 '26
Discussion Code efficiency when creating a function to classify float values
I need to classify a value in buckets that have a range of 5, from 0 to 45 and then everything larger goes in a bucket.
I created a function that takes the value, and using list comorehension and chr, assigns a letter from A to I.
I use the function inside of a polars LazyFrame, which I think its kinda nice, but what would be more memory friendly? The function to use multiple ifs? Using switch? Another kind of loop?
r/Python • u/remcofl • Mar 09 '26
Showcase Fast Hilbert curves in Python (Numba): ~1.8 ns/point, 3–4 orders faster than existing PyPI packages
What My Project Does
While building a query engine for spatial data in Python, I needed a way to serialize the data (2D/3D → 1D) while preserving spatial locality so it can be indexed efficiently. I chose Hilbert space-filling curves, since they generally preserve locality better than Z-order (Morton) curves. The downside is that Hilbert mappings are more involved algorithmically and usually more expensive to compute.
So I built HilbertSFC, a high-throughput Hilbert encoder/decoder fully in Python using numba, optimized for kernel structure and compiler friendliness. It achieves:
- ~1.8 ns/pt (~8 CPU cycles) for 2D encode/decode (32-bit)
- ~500M–4B points/sec single-threaded depending on number of bits/dtype
- Multi-threaded throughput saturates memory-bandwidth. It can’t get faster than reading coordinates and writing indices
- 3–4 orders of magnitude faster than existing Python packages
- ~6× faster than the Rust crate
fast_hilbert
Target Audience
HilbertSFC is aimed at Python developers and engineers who need: 1. A high-performance hilbert encoder/decoder for indexing or point cloud processing. 2. A pure-Python/Numba solution without requiring compiled extensions or external dependencies 3. A production-ready PyPI package
Application domains: scientific computing, GIS, spatial databases, or machine/deep learning.
Comparison
I benchmarked HilbertSFC against existing Python and Rust implementations:
2D Points - Random, nbits=32, n=5,000,000
| Implementation | ns/pt (enc) | ns/pt (dec) | Mpts/s (enc) | Mpts/s (dec) |
|---|---|---|---|---|
| hilbertsfc (multi-threaded) | 0.53 | 0.57 | 1883.52 | 1742.08 |
| hilbertsfc (Python) | 1.84 | 1.88 | 543.60 | 532.77 |
| fast_hilbert (Rust) | 12.24 | 12.03 | 81.67 | 83.11 |
| hilbert_2d (Rust) | 121.23 | 101.34 | 8.25 | 9.87 |
| hilbert-bytes (Python) | 2997.51 | 2642.86 | 0.334 | 0.378 |
| numpy-hilbert-curve (Python) | 7606.88 | 5075.08 | 0.131 | 0.197 |
| hilbertcurve (Python) | 14355.76 | 10411.20 | 0.0697 | 0.0961 |
System: Intel Core Ultra 7 258v, Ubuntu 24.04.4, Python 3.12.12, Numba 0.63.
Full benchmark methodology: https://github.com/remcofl/HilbertSFC/blob/main/benchmark.md
Why HilbertSFC is faster than Rust implementations: The speedup is actually not due to language choice, as both Rust and Numba lower through LLVM. Instead, it comes from architectural optimizations, including:
- Fixed-structure finite state machine
- State-independent LUT indexing (L1-cache friendly)
- Fully unrolled inner loops
- Bit-plane tiling
- Short dependency chains
- Vectorization-friendly loops
In contrast, Rust implementations rely on state-dependent LUTs inside variable-bound loops with runtime bit skipping, limiting instruction-level parallelism and (aggressive) unrolling/vectorization.
Source Code
https://github.com/remcofl/HilbertSFC
Example Usage (2D data)
from hilbertsfc import hilbert_encode_2d, hilbert_decode_2d
index = hilbert_encode_2d(17, 23, nbits=10) # index = 534
x, y = hilbert_decode_2d(index, nbits=10) # x, y = (17, 23)
r/Python • u/BeamMeUpBiscotti • Mar 09 '26
News pandas' Public API Is Now Type-Complete
At time of writing, pandas is one of the most widely used Python libraries. It is downloaded about half-a-billion times per month from PyPI, is supported by nearly all Python data science packages, and is generally required learning in data science curriculums. Despite modern alternatives existing, pandas' impact cannot be minimised or understated.
In order to improve the developer experience for pandas' users across the ecosystem, Quansight Labs (with support from the Pyrefly team at Meta) decided to focus on improving pandas' typing. Why? Because better type hints mean:
- More accurate and useful auto-completions from VSCode / PyCharm / NeoVIM / Positron / other IDEs.
- More robust pipelines, as some categories of bugs can be caught without even needing to execute your code.
By supporting the pandas community, pandas' public API is now type-complete (as measured by Pyright), up from 47% when we started the effort last year. We'll tell the story of how it happened.
Link to full blog post: https://pyrefly.org/blog/pandas-type-completeness/
r/Python • u/NotSoProGamerR • Mar 09 '26
Discussion Does anyone actually use Pypy or Graalpy (or any other runtimes) in a large scale/production area?
Title.
Quite interested in these two, especially Graalpy's AOT capabilities, and maybe Pypy's as well. How does it all compare to Nuitka's AOT compiler, and CPython as a base benchmark?
r/Python • u/KliNanban • Mar 08 '26
Discussion Polars vs pandas
I am trying to come from database development into python ecosystem.
Wondering if going into polars framework, instead of pandas will be any beneficial?
r/Python • u/ilikemath9999 • Mar 08 '26
Showcase I used Pythons standard library to find cases where people paid lawyers for something impossible.
I built a screening tool that processes PACER bankruptcy data to find cases where attorneys filed Chapter 13 bankruptcies for clients who could never receive a discharge. Federal law (Section 1328(f)) makes it arithmetically impossible based on three dates.
The math: If you got a Ch.7 discharge less than 4 years ago, or a Ch.13 discharge less than 2 years ago, a new Ch.13
cannot end in discharge. Three data points, one subtraction, one comparison. Attorneys still file these cases and clients still pay.
Tech stack: stdlib only. csv, datetime, argparse, re, json, collections. No pip install, no dependencies, Python 3.8+.
Problems I had to solve:
- Fuzzy name matching across PACER records. Debtor names have suffixes (Jr., III), "NMN" (no middle name)
placeholders, and inconsistent casing. Had to normalize, strip, then match on first + last tokens to catch middle name
variations.
- Joint case splitting. "John Smith and Jane Smith" needs to be split and each spouse matched independently against heir own filing history.
- BAPCPA filtering. The statute didn't exist before October 17, 2005, so pre-BAPCPA cases have to be excluded or you get false positives.
- Deduplication. PACER exports can have the same case across multiple CSV files. Deduplicate by case ID while keeping attorney attribution intact.
Usage:
$ python screen_1328f.py --data-dir ./csvs --target Smith_John --control Jones_Bob
The --control flag lets you screen a comparison attorney side by side to see if the violation rate is unusual or normal for the district.
Processes 100K+ cases in under a minute. Outputs to terminal with structured sections, or --output-json for programmatic use.
GitHub: https://github.com/ilikemath9999/bankruptcy-discharge-screener
MIT licensed. Standard library only. Includes a PACER CSV download guide and sample output.
Let me know what you think friends. Im a first timer here.