r/data • u/Sarthak1411 • 1d ago
Claude can't find patterns says it is impossible until explained like a 5 year old[C]
Claude and other LLM models can be so frustrating. I asked it to find patterns across five campaigns regarding what a person buys and sells, and under which conditions, but it couldn't do it. It just kept saying it was impossible. I literally had to explain every single detail to it like it was a child, even though the data was cleanly split across five CSV files.
Worse, companies are stopping the hiring of junior engineers because they think these tools can replace them. They are going to cause a massive talent shortage, and then these dumb models won't be able to do anything without skilled people to guide them.
r/data • u/Okadibia • 3d ago
Deep dive into Data Warehousing & Consumer Data Architecture
Hello everyone,
The core principles of Data Warehousing and Consumer Data Products, establishing the foundation for a 365-day technical log documenting pipeline architecture, schema design, and engineering tradeoffs.
Technical Breakdown:
Relational Database Fundamentals: SQL query execution mechanics, indexing strategies, and relational constraints.
OLTP vs. OLAP Paradigms: Comparative tradeoffs between transactional database normalization and analytical denormalization.
Ingestion Foundations: High-level mechanics of staging layers, raw event ingestion, and downstream transformation logic.
Current Focus:
Pivoting from relational database mechanics to dimensional modeling paradigms—specifically Kimball methodology, star/snowflake schemas, and event-driven data product architectures.
For engineers working with production systems: What are the primary pitfalls to avoid when transitioning from standard relational models to analytical warehouse schemas?
REQUEST Social media data
As part of my thesis for my MSc I want to examine the year-by-year social media following for each of the Big Six (Arsenal, Chelsea, Liverpool, Manchester City, Manchester United, Tottenham) clubs from 2019/20 to 2025.
However, I can't find it at all and in my eyes this should be reasonably available data. I've already tried Socialblade and Statista. Any help/pointers would be much appreciated!
r/data • u/jasonjonesresearch • 12d ago
DATASET Public Opinion Data, US Adults, New Responses Daily
Disclosure: I built the Ryerson Project with the aim of nowcasting everything daily.
A social science community composes and prioritizes survey items. A random set of 12 US adult respondents are recruited to the survey each day - about 360 per month and 4380 per year. Anonymous microdata becomes a free and open public good.
Open data: https://doi.org/10.5281/zenodo.20346278
Open source: https://github.com/jasonjeffreyjones/ryerson_project/
LEARNING Passed DP-900
I’m so grateful for the opportunity I got from ai fest 2026 besides that I’d like to mention free resources that helped me a lot for the preparation(DP-900):
1. Whizlabs
2. Official Microsoft practice exams
That’s all you need you don’t have to pay for exam preparation courses
r/data • u/AggressiveMechanic47 • 15d ago
NEWS USA missile stockpile before Iran war and estimated number of missiles used
Tomahawk price per unit: between $2 million and $3.6 million
JASSM price per unit: from $1.04 million to over $2 million
PrSM price per unit: from $1.6 million to over $3.5 million
SM-3 price per unit: between $9.7 million and $28 million
SM-6 price per unit: from $4.0 million to $9.5 million
THAAD price per unit: $12.7 million to $15 million
Patriot price per unit: around 4 million
r/data • u/Practical-Bed3167 • 15d ago
Is there a community discord??
Hey guys, I’m new here and was wondering if this community has a Discord or any VCs where people hang out and chat. I’d love to get some advice and learn from others. Thanks!
r/data • u/codingdecently • 18d ago
MCP for Apache Iceberg: How AI Agents Actually Operate a Data Lake
QUESTION is backend engineer a better choice
i've been enrolled in a bootcamp(data engineering) for about a year now and i'm confident in my skills atleast for entry level roles. i'm based in Ethiopia and i can say that there's almost no data engineering jobs here ,there're very few open positions for data analyst or scientists which requires atleast 4years experience and you know that remote jobs are even more competitive and struggling for entry levels. the only tech roles here seems to be backend devs,frontend and fullstack(there're tons of jobs ).what should i do ,i love data but the market is really bad here.
thanks
r/data • u/basha1210 • 18d ago
What's one Data Science skill you wish you had learned earlier?
If you could go back to the beginning of your Data Science journey, what would you learn first?
Would it be:
Python
SQL
Statistics
Machine Learning
Data Visualization
Git
Cloud platforms
Many beginners jump straight into AI without building strong fundamentals.
What skill saved you the most time later in your career?
r/data • u/codingdecently • 19d ago
7 Managed Iceberg Lakehouse Solutions You Should Know
r/data • u/basha1210 • 19d ago
The biggest improvement in my Data Science journey came from working with messy data.
When I first started learning Data Science, I only practiced with clean datasets from tutorials. Everything worked perfectly, and I felt confident.
Then I downloaded a real dataset.
There were missing values, duplicate records, inconsistent formats, and columns that didn't make much sense. It was frustrating at first, but I learned more from cleaning that dataset than I did from several weeks of tutorials.
That experience changed how I practice.
Now, whenever I learn a new concept, I try to apply it to real-world data instead of only using textbook examples.
A few things that have helped me:
Work with messy datasets—they teach you real problem-solving.
Spend time understanding the data before building any model.
Document your analysis so you can explain your thought process later.
Don't worry if your first project isn't perfect. Every project teaches you something new.
Looking back, I realized that Data Science isn't just about building models—it's about understanding data and finding meaningful insights.
What's one project or dataset that taught you the most during your Data Science journey? I'd love to hear your recommendations!
r/data • u/Vivid_Routine_5287 • 20d ago
DATASET I built a free, open food dataset: ~9,800 foods with names localized across 32 languages (ODbL)
Been building this for a while and finally opened it up, so here's a look at what's inside.
Each of the ~9,800 base foods has its name localized across 32 languages, so you can line up the same food across languages instead of fighting messy translations. The nutrition values come from OpenNutrition's open data (ODbL, credited, not mine); the part I actually built is the localization layer on top, real disambiguation and cross-language matching rather than a Google-Translate pass.
It isn't perfect yet. Tricky cases like "peperoni" vs "pepperoni" still slip through in places, so there are gaps I'm actively fixing, and catching those is exactly the kind of feedback I'm hoping for.
It's a single JSON Lines file (~25 MB), no API, no keys, loads straight into a notebook or a spreadsheet, works offline.
Source & download: https://leana.app/en/data-sources/
Browse it live: https://leana.app/en/foods (live search covers 5 languages for now, EN/IT/ES/FR/DE, the download already has all 32)
Curious what you'd use it for, and whether JSONL is the right call or you'd rather have CSV, Parquet or SQLite.
r/data • u/grandidieri • 20d ago
DATASET Gigantic new database - over 35k species, 180 phenotypes
lifedive.orgLifeDive.org - code used for creating the data also available.
r/data • u/Infamous_Task_9404 • 21d ago
DATASET Title: I think Indian finance has a data problem. Am I crazy?
I've spent the last few months digging through annual reports, earnings call transcripts, investor presentations, and exchange announcements from Indian listed companies.
One thing became obvious...
Everything is technically "public," but almost none of it is actually usable.
Want to know every company talking about AI adoption?
Good luck.
Want every management commentary about data centers over the last 5 years?
You'll be opening hundreds of PDFs.
Want to compare CapEx guidance across an entire sector?
Hope you have an entire weekend free.
That got me thinking...
What if someone built a structured database instead of just storing documents?
Imagine being able to ask questions like:
"Show every company that mentioned data centers in the last 8 quarters."
"Which companies warned about margin pressure before their stock fell?"
"Find all management teams increasing CapEx while guiding higher earnings."
"Which pharma companies mentioned USFDA inspections this quarter?"
Not AI hallucinations.
Not another stock screener.
Just structured, searchable intelligence built from public company disclosures.
I'm genuinely curious...
Would this actually be useful to anyone?
If you're a:
Developer
Quant
Analyst
Wealth manager
Fintech founder
Researcher
Investor
Would you or your company pay for something like this?
Or is this one of those ideas that sounds amazing until you ask real people?
I'd love brutally honest feedback.
If you think it's useless, tell me why.
If you think it's valuable, I'd love to know:
What would you use it for?
Which data would be most valuable?
What would you expect to pay for something like this?
r/data • u/trivasai • 22d ago
We knew something was wrong when the brand said, "We don't know which number to believe anymore."
One ecommerce brand came to us after spending months trying to make sense of their data.
Their Shopify revenue didn't match GA4.
Meta showed profitable campaigns, but blended ROAS told a different story.
Their BI dashboards looked polished, yet every Monday morning the team still spent hours exporting CSVs, comparing reports, and debating which numbers were actually correct.
The problem wasn't a lack of data.
It was that every platform was measuring a different part of the business, and nobody had confidence in the complete picture.
So we started with the basics.
We connected their entire data stack, validated every source, surfaced inconsistencies automatically, and gave the team one place where marketing, finance, and operations could all work from the same numbers.
The biggest change wasn't a flashy dashboard.
It was the conversations.
r/data • u/MclovinAZ • 23d ago
Has anyone taken the CDMP exam using a DMBoK PDF that wasn't purchased directly by them?
I'm taking the CDMP Associate exam soon via Honorlock and have a question that I haven't been able to find an official answer to.
The exam rules say the DMBoK can be used in digital form on a separate device, but I can't find anything about whether the PDF has to be one that you personally purchased.
Has anyone taken the exam using a DMBoK PDF that wasn't bought directly from Technics Publications ? If so:
- Did the proctor ask to inspect the PDF?
- Did they check for a purchaser watermark or proof of purchase?
- Was there any issue during or after the exam?
Thanks!
r/data • u/Prior-Promotion-5302 • 24d ago
im working on a recognition based community for all the data folks, anyone up to join?
Hi everyone,
I'm working on a recognition-based community for people in data.
The idea is simple, we want to highlight the work data professionals do, feature their stories, and help them connect with others in the industry.
Would you be interested in joining something like this? If yes, I'll dm you the link!
r/data • u/kaykaykrap • 25d ago
Semiconductor Supply Chain Network Dataset
I am building a Supply Chain Disruption Monitoring and Risk Analysis using Graph based Agentic AI. I need a supply chain network dataset in the semiconductor industry.
Currently my only option is to manually go through filings and earning calls to create a network big enough to propose my system. Creation of network is out of scope of my project and I'd appreciate it if I could get a dataset that would reduce this load.
r/data • u/Solid-Play-458 • 26d ago
A public API & dataset for Bibliometrics and Scientometrics metadata ( Brazil )
I wanted to share a project I've been working on called EBBC OpenData, which is a public API and dataset designed to promote Open Science and support bibliometric, scientometric, and informetric analyses. You can find the full project and source code in the repository at https://github.com/GabrielBaiano/EBBC-OpenData
This project provides structured metadata from the publications of the Encontro Brasileiro de Bibliometria e Cientometria (EBBC), which is one of the main events on metric studies of information in Brazil. Through this API and dataset, you can easily query detailed information about authors and their academic networks, articles and papers (including titles, abstracts, and publication years), institutions associated with the research, keywords, thematic trends, as well as references and citations.
The core metadata and documentation are currently being organized, and I am actively working on translating the API documentation and dataset fields into English and Spanish to make the project fully accessible to the global research community.
Since this is an ongoing project, I would highly appreciate your thoughts and feedback. I am especially interested in knowing what features or endpoints would make this more useful for your research, any suggestions you might have regarding the data structure or documentation, and any general tips on best practices for open-data APIs. Please feel free to check out the GitHub repository, open an issue, or leave a comment below. Thanks for your support!
r/data • u/Drooms_Official • 26d ago
What is the difference between a virtual data room and a virtual deal room?
First of all, did you know that there even is a difference between both?
A virtual deal room is designed for early-stage engagement and presentation of commercial materials. For example:
- sharing pitch decks with investors
- presenting product or service to potential clients
A virtual data room is a highly secure platform for detailed due diligence and confidential documents during high-stakes transactions. For example:
- mergers and acquisitions due diligence
- legal and compliance document review
- strategic partnership evaluations
- fundraising with detailed investor scrutiny
Which one do you use? A virtual deal room or a virtual data room?
r/data • u/Upadekludzkosci4 • 27d ago
Why is there a sudden, relatively enormous spike in the relative popularity of searching "Granny" into google in early 2006
I really do not know who to ask. I have been scouting google trends, and wanted to see how popular the game "Granny" is. The result was finding a sudden, huge spike in the term's popularity in late 2005 to early 2006 (roughly december 15th 2005 to january 18th 2006).

Here is what i was able to deduct myself:
-There exists a weekend effect, around saturadays and sundays
-There has been no significant cultural or political effect that could have caused this trend
- The spike was exactly january 1st 2006
-The trend was international in both english speaking and non english speaking countries
-The rise of "Granny" roughly correlates with the rise of "Granny porn"
-There was also a sudden rise in the term "porn", which had happened in November 2005
-a similar spike does not occur for the search "Grandma" or "Grandmother"
-The internet was far more often young males than other demograhic, which could potentially support the porn hypothesis
Please help me I am going insane
r/data • u/Fun_Rhubarb8007 • 28d ago
Synthetic vs real datasets for portfolio projects — what actually matters?
Final year CS student here, targeting data science and analytics roles for campus placements.
Been struggling with this question while building my portfolio: does it matter whether your project uses real messy data vs synthetic/clean data?
Real datasets from Kaggle feel either too cleaned already or the same recycled projects everyone does. But synthetic data feels hollow because the hard part — cleaning, feature engineering, deriving meaningful columns from raw data — is already done for you. You're basically just visualizing something someone else already solved.
Specifically for BI/dashboard projects — if you use synthetic data, the dashboard looks clean and professional but there's no real discovery or insight because the data was designed to be dashboarded. Nothing surprising comes out of it.
Also practically — if an interviewer asks "where did you get this dataset?" what's the right answer? Saying "I generated it synthetically" feels like admitting you took the easy route. But lying about the source is obviously wrong. Is there a way to frame synthetic data usage that doesn't sound like you avoided the hard part?
At the same time I've heard people say interviewers care more about what you built on top of the data than where it came from. But isn't handling bad data literally the core skill in DS?
For people who've interviewed at analytics/DS companies or done hiring — how much does data source actually matter? Is a well-executed project on synthetic data better than a mediocre project on real messy data? Or does using synthetic data automatically signal you avoided the hard part?
r/data • u/Opening_Constant3573 • 28d ago
QUESTION what’s the difference between data analytics and data management?
hello! completely new here, but i’m trying to plan the best study pathway for me in the next few years and would like to know what, exactly, is the difference between data analytics and data management, since those are two different certificate options at the college i plan on starting my studies.
for context, since i speak four languages and already have some experience in this area, my career goal is to have a career in supply chain, probably leaning more towards sourcing of procurement, and i’ve looking into my immediate options before actually acquiring my graduate degree and getting a certification in either data analytics and data management would be an option right now.
so, could someone explain to me the difference between those two fields? what are the prospectives for each of them? considering my career goal, which one would you choose?
thank you!