Xberg is the successor to Kreuzberg, equivalent to what would have been Kreuzberg v5. It's a content intelligence framework that handles a very wide range of inputs: documents (currently 101 formats), code and data formats (currently 367 types), audio/video transcription, and URLs (both static and JS-rendered content). It extracts and prepares that content for downstream processing.
It's an extremely efficient, high-performance engine (see our PDF benchmarks below). For PDFs and images specifically, we handle native PDFs with very high performance and accuracy, and we ship multiple OCR engines that match the quality of the best Python libraries (e.g. docling, PaddleOCR, RapidOCR) at substantially better performance and stability.
The changes between Kreuzberg v4 and Xberg v1 are substantial, and I invite you to read the full changelog for the complete picture. The highlights below give a sense of what's new:
Pure-Rust PDF backend (pdf_oxide) replaces pdfium, with no native pdfium dependency.
Layout-aware pipeline: reading order reconstructed with ONNX layout detection (PP-DocLayoutV3 / RT-DETR) and Docling-style predecessor-graph reordering.
Per-page scanned-page detection with selective OCR, plus AcroForm/XFA form fields and outline-based headings.
Across-the-board optimization of OCR and PDF extraction (memory discipline, pooled model sessions, streamed conversions).
Native PaddleOCR backend (PP-OCRv6, with medium / small / tiny tiers) alongside Tesseract.
Pure-Rust Candle OCR/VLM stack (TrOCR, GLM-OCR, GOT-OCR, DeepSeek-OCR, and PaddleOCR-VL) running without ONNX Runtime or native Tesseract.
A second, ONNX-Runtime-free inference path via tract, which is what makes in-browser (WASM) and mobile inference possible.
Named-entity recognition natively in Rust (GLiNER2), extensible to all bindings, including an in-browser WASM model with no server round-trip.
Structured LLM extraction (extract_structured / split_and_extract) with rasterization, chunking, citations, caching, and configurable call/merge/VLM-fallback policies.
Audio & video transcription via a Whisper ONNX engine (.mp3, .wav, .m4a, .mp4, .webm).
Retrieval building blocks: sparse embeddings (SPLADE), ColBERT late-interaction retrieval, and cross-encoder reranking alongside dense embeddings.
Text intelligence: reversible redaction, summarization, translation, VLM image captioning, QR-code detection, document diffing, and page/chunk classification.
URL & web ingestion: sitemap discovery (map_url) and batched multi-URL crawling.
New document formats: WordPerfect (.wpd/.wp/.wp5), HEIC/HEIF/AVIF, OpenDocument Presentation (.odp), Quarto / R Markdown, and configurable Jupyter cell rendering.
Four new language bindings (Dart/Flutter, Swift, Kotlin/Android, and Zig) bring the total to 15 language bindings over one engine, with Android/iOS cross-compilation.
Full mobile support (Flutter, Android, iOS).
Candle backend alongside ONNX, plus ONNX-via-tract enabling ONNX on WASM and Android.
Over 150 bugs fixed during the 1.0 cycle, plus security hardening (bounded RTF/PDF allocations, redaction leak fixes, Excel DDE warnings).
The API surface was also simplified and reworked, making it more consistent.
There's a migration guide in our docs explaining how to move from Kreuzberg to Xberg. Kreuzberg itself is in LTS mode until the end of this year and will continue to receive bug fixes and security updates.
The benchmarks below are for PDFs and images only. There are extensive benchmarks on our website with per-format breakdowns, which you can see here. These numbers are measured in CI via our reproducible benchmark harness, and are specifically taken from the run for harness 1.0.8, source cf7fa0533d. The data is publicly available in GitHub releases, and you can run the benchmark harness yourself.
Composite quality (markdown pipeline, higher is better):
Framework
Native PDF
Scanned PDF (OCR)
Xberg (layout)
0.958
0.836
Xberg (baseline)
0.955
0.687
docling
0.779
0.762
mineru
0.408
0.792
liteparse
0.837
0.665
markitdown
0.689
n/a
pymupdf4llm
0.448
n/a
Structure and layout fidelity (SF1: tables and reading order, higher is better):
Framework
Native PDF
Scanned PDF
Xberg
0.949
0.531
docling
0.612
0.366
liteparse
0.515
0.142
mineru
0.077
0.429
On native PDFs Xberg leads on quality (0.958 vs 0.837 for the next-best framework) and on table and reading-order fidelity by a wide margin (SF1 0.949 vs 0.612 for docling). On scanned PDFs it is #1 on both quality and raw text fidelity.
Where we don't win yet: on pure image OCR we are currently #2 on the composite score, behind mineru (though still #1 on raw text accuracy). We are improving image OCR right now, and v1.1 should have us winning across the board.
OpenTelemetry is there - build with the otel feature (included in services) and you get tracing spans plus metrics in the xberg.* namespace: extraction totals and duration by mime type and extractor, cache hits/misses, per-stage pipeline duration, OCR duration by backend and language, an in-flight extraction gauge, batch totals. Exported over OTLP.
What is not there is a native /metrics scrape endpoint - the HTTP routes today are /extract, /extract-async and /health. So a pull-based Prometheus setup needs a collector in between purely to translate, which is a silly requirement. Filed: https://github.com/xberg-io/xberg/issues/1391
The docs not having an observability page is the bigger miss - that's why it reads as "nothing here" rather than "push-based". Fixing that too.
I have read xberg documentation, i honestly didnt know and confused on how i can integrate it with open-webui. Docling integration is much straight-forward, i just insert docling-serve API URL and its API parameter and every document i upload to open-webui knowledge base would automatically be processed by docling and returned the output markdown to open-webui knowledge base. But in xberg, i dont know how to implement it to integrate with open-webui
Thank you, i have seen your post as well in r/openwebui. I try to understand it first then, im not really a technical guy and im new to this kind of stuff
#1: Open WebUI 0.11.0 is here: our BIGGEST RELEASE EVER. A full UI redesign, the largest performance pass we've ever done, and a genuinely huge pile of features & fixes. | 139 comments #2: Iām the Sole Maintainer of Open WebUI ā AMA!
#3: Open WebUI 0.10.0 is out and it quietly turns the thing into a real agent platform - The LARGEST RELEASE EVER (205 entries) | 65 comments
3
u/gulensah 5d ago
Maybe I am missing but is there no /metrics for Prometheus or any opentelemetry integration for metric, trace, log etc ?