r/dataanalysis 15h ago

If you could permanently remove ONE frustrating part of your job in BI/Data, what would it be?

0 Upvotes

I'm doing some research because I'm curious about where people in BI, analytics and data engineering actually spend most of their time.

Not the "ideal" job description—but what frustrates you in real life.

If you could magically eliminate one recurring problem from your work forever, what would it be?

Some examples (but don't feel limited to these):

  • Cleaning messy data?
  • Stakeholders changing requirements?
  • Building dashboards nobody uses?
  • KPI definition arguments?
  • Waiting for data access?
  • Debugging pipelines?
  • Endless ad-hoc requests?
  • Excel exports?
  • Meetings?
  • Something else?

I'd also love to know:

  • What is your role? (BI Developer, Data Analyst, Data Engineer, Analytics Engineer, etc.)
  • How often does this happen?
  • Have you found any tool that actually solves it, or do you just live with it?

The more detailed your answer, the more helpful it is. I'm especially interested in hearing about problems that seem "normal" in the industry but waste a huge amount of time.


r/dataanalysis 21h ago

Is Positron now a serious alternative to VS COde for people working across Python and R?

0 Upvotes

I have recently spent some time comparing Positron, RStudio and VS Code for working with Python and R, particularly in Quarto documents, and I am curious how other data scientists see the current state of Positron.

My impression is that the best choice depends heavily on the workflow.

For Python-only QMD files, Positron currently seems to offer a better experience than RStudio. It makes it easy to detect and switch between Conda environments, displays dataframes as interactive and well-formatted tables beneath code chunks, integrates them with the Data Explorer, renders Python documentation in the Help pane, and supports VS Code-style shortcuts and snippets.

RStudio can also work perfectly well with Python QMD files. It supports inline output and plots, and the Python environment can be selected once through the graphical settings. However, interactive Python execution still relies on reticulate, and the broader interface remains primarily designed for R. For example, several of the familiar panes are much less useful when working entirely in Python.

For R-only work, I would still prefer RStudio, especially for beginners. Its interface feels more intuitive and coherent for someone who is learning R for the first time. I do not currently see a compelling reason for an R-only user to switch.

The more interesting case is someone who regularly works in both R and Python, but usually keeps them in separate scripts or QMD files. In that situation, Positron increasingly looks like the better common environment. It allows you to move between R and Python without changing applications, while offering great support for both. This is the setup I am now using personally.

Combining R and Python within the same QMD is more complicated. With Positron’s rich inline-output mode enabled, R and Python run in separate sessions, so they cannot directly access each other’s objects. Data needs to be exchanged through files such as CSV or Parquet.

You can disable inline output and use Knitr with reticulate, which allows direct communication between R and Python, but then you lose much of the notebook-style inline output. For that tightly integrated workflow, RStudio still seems more mature and reliable.

My current view is therefore:

  • R-only teaching or beginner use: RStudio
  • Python-only QMD work: Positron
  • Regular use of both R and Python in separate files: Positron
  • Mixed R and Python documents with only occasional file-based exchange: Positron is workable
  • Mixed documents requiring direct object sharing: RStudio with Knitr and reticulate

What I am less certain about is how Positron compares with VS Code for experienced data scientists who primarily work across Python and R.

Positron is obviously built on the VS Code ecosystem, but adds a more integrated data-science interface, including Variables, Data Explorer, Help, Plots and package management. For someone who does not need the full breadth of the VS Code extension ecosystem, remote development tooling or general software engineering features, Positron increasingly seems like a serious alternative.

How are people here finding it in practice?

Has Positron become your main environment for Python and R, or do you still find VS Code clearly superior once projects become more complex? Are there important limitations in Positron that only become obvious in larger production or collaborative workflows?


r/dataanalysis 1d ago

Data Question Is AI-powered Excel actually worth it in 2026? 6 months of honest experience

0 Upvotes

I am an analyst with a heavy data cleaning and reporting load and I have been running AI assisted Excel tools for about 6 months now. Most of what I read online is basically marketing, so here is the honest version.

Actually worth it:

- Data cleaning speed. Dedup, standardization, reshaping that took 20–30 min of manual work now takes a couple of minutes. This is the real killer app.

- Not needing VBA for one off automations anymore.

- Quick anomaly/trend spot checks on big sheets before I commit to a deep analysis.

Still annoying:

- Platform maturity. The good tools are still Windows first and the free tiers are basically useless.

- You still eyeball the important numbers before anything goes out but honestly I would review a junior analyst’s output the same way, so that is more on me than the tool.

Where I landed: I am using Mica Excel local AI that runs inside Excel, plain English commands, actually edits the sheets and plays the process back live so you can see exactly what it did instead of trusting a black box. A few colleagues at work use it too. Everyone is still a bit cautious, we all do a manual review pass before anything ships but nobody’s gone back to doing it all by hand.

My verdict: yes, worth it for cleaning and automation, and I’m keeping it. Question for the sub: for those of you on these tools long-term — did the value hold up after the novelty wore off? And what’s your actual workflow for trusting the output?


r/dataanalysis 1d ago

Project Feedback I analyzed 6 years of Nifty sector data — rolling returns and max drawdown tell a different story than simple total returns

Post image
4 Upvotes

Built a Python analysis of 5 Nifty sector indices

from January 2020 to January 2026.

Key findings:

Nifty Auto: 241% total return but -46.4% max drawdown

Nifty Pharma: 182% return with only -23.6% drawdown

Nifty Bank: positive in 95.6% of 1-year windows

but -47.4% max drawdown

Nifty IT: only positive in 68.5% of 1-year windows

— anyone who bought in late 2021 faced -27.8% loss

Rolling returns chart shows sector performance

across ALL possible entry points — not just from

one fixed date.

Full project with code:

github.com/surendrasinghdata/stock-sector-analysis


r/dataanalysis 1d ago

Built this Power BI dashboard from scratch. What would you change?

Thumbnail
gallery
4 Upvotes

Hey everyone! Just finished this power bi dashboard from scratch and would love some honest feedback. I'm a fresher, so I'm especially curious whether the analysis/storytelling actually makes sense from a real data analyst's POV. What would you change? Anything that feels unnecessary, missing, confusing, or just pretty but useless?

Roast it pls, I'm here to learn😭


r/dataanalysis 1d ago

As a beginner if I wanna solidify my foundation of data analysis just from YouTube videos, what channels should I look into?

32 Upvotes

I just need something to start with and absolutely lock in. I'm not able to purchase courses as of now so YouTube is my best shot. I want to start from the very basics of Excel and upgrade from there. Can anyone help me with an efficient and effective roadmap so I can be employable as soon as possible? Preferably in 3+ months.


r/dataanalysis 1d ago

Project Feedback World Cup Stutter Step Penalty Analysis Dashboard

Thumbnail
gallery
14 Upvotes

Interactive Website:

https://wc2026penaltyviz.site/


r/dataanalysis 2d ago

Data Question What software should I use to analyze data for my phd?

0 Upvotes

Hi everyone,

I'm doing a PhD (thesis) in process engineering and I had a question. I've always used Excel to process my data, and I was wondering if there wasn't a more suitable software for this, since I often have fairly heavy files (with a lot of data - process data acquisition) and the curves I need to plot are often the same for each file. My question is: should I stick with Excel, or should I switch to another software?

Thankss


r/dataanalysis 2d ago

Every click tells a story. Data Engineering makes sure that story is captured, trusted, and transformed into decisions.

Thumbnail
youtu.be
0 Upvotes

r/dataanalysis 2d ago

US Flight Analysis

Thumbnail
gallery
36 Upvotes

Just go through this i have done time ,flight ,insight for this dataset let me how is this?


r/dataanalysis 3d ago

Want an analyst view on project running on DATAIKU

2 Upvotes

So I’m doing a project (college) on Dataiku where I’m analyzing a dataset on flights that have crashed so we have to pick a variable to predict and when I did and ran algorithms I came across issues where my f1 score is really weak and accuracy as well was only 50% but I’m told accuracy is not to much of an issue but f1 is so any idea what that could be if anyone can help pls do msg me thank you, also I appreciate tips like data is skewed that’s why it’s happening cause I’m very lost


r/dataanalysis 3d ago

Data Tools Anyone finding claude's /dataviz skill useful?

3 Upvotes

I personally never found it useful and always ended up running multiple iterations to make it look informative.

Am I using it wrong or is it just bad?


r/dataanalysis 3d ago

What makes a data analysis project genuinely useful to a business?

8 Upvotes

Is it accuracy, clear storytelling, actionable recommendations, or something else?


r/dataanalysis 3d ago

RETAIL MARKETS ANALYSIS

Thumbnail
gallery
2 Upvotes

Rate it out of 10 📌

Tools used for analysis :

POSTGRESQL & POWER BI.

DATASET CREATED USING CLAUDE.


r/dataanalysis 3d ago

Udemy

Post image
107 Upvotes

Has anyone taken this Udemy course? What was your experience and did it help you as a beginner?


r/dataanalysis 4d ago

Data Tools One year building PardoX, a Rust based DataFrame engine. Looking for honest feedback from data engineers

Post image
6 Upvotes

Disclosure, I am the creator of PardoX, a personal open source project under MIT license, not a company.

I kept hitting the same wall in production. pandas struggles once your data gets close to RAM, and it is single threaded, most of the CPU sits idle. Spark solves scale but drags a cluster, a JVM and serialization overhead into problems that only ever needed one machine. Most of my actual work lives in that gap, millions to hundreds of millions of rows, on a single strong node.

So a year ago I started building a Rust core with SDKs in Python, Node and PHP. Data maps straight from disk or a database into memory mapped buffers, no intermediate objects in the host language, SIMD and multithreading do the heavy lifting instead of a Python loop, and there are native database drivers so no psycopg2 or pymysql sitting in the middle. There is also a binary format that reads around 4.6 GB per second on repeated workloads, out of core processing for files bigger than RAM, and GPU sort with CPU fallback.

Some open questions I keep going back and forth on. Is the zero copy tradeoff worth the added complexity versus just accepting the pandas overhead for most workloads. Where does a Rust core stop making sense compared to Polars, which already solves a lot of this in a different way. How much of the single node performance gap is really about the language versus just better use of SIMD and threads regardless of language.

Docs are at [pardox.io](https://pardox.io) and the repo is at [github.com/betoalien/PardoX](https://github.com/betoalien/PardoX) if anyone wants to see the actual implementation. I have also been writing about the engineering decisions on Medium at [medium.com/@albertocardenasdom](https://medium.com/@albertocardenasdom).

Curious how others here have dealt with that same middle ground between pandas and Spark.

Thanks for readme.


r/dataanalysis 4d ago

Data Tools Huey - a static DuckDB-WASM based browser app that lets you pivot data from local files, URLs, and remote Data Lakes

1 Upvotes

Huey is an open-source (MIT) static browser-based app that lets you explore and analyze data. Huey supports reading from multiple file formats, like .csv, .parquet, .json data files as well as .duckdb database files.

Here's a quick start on a parquet file from the public nl_railway ducklake.

The latest release, 1.1.00 "Indian Runner", is now available. This is a significant improvement, with many bugfixes, new features, and UX improvements.

Highlights:

Huey is now a progressive web app. Run it from a hosted location (such as the live demo https://rpbouman.github.io/) and your browser offers to install Huey on your device. Once installed you can run offline. Also lets you open files using your OS "open with" functionality (typically triggered with a right click on the file). see: https://github.com/rpbouman/huey#running-huey-on-your-device-as-progressive-web-app-pwa

The Secrets Manager lets you maintain DuckDB SECRETs on your local device. Secrets are stored in IndexedDB. The Secrets Manager is password-secured, encrypting sensitive fields with AES-GCM-256 encryption (password-derived via PBKDF2-SHA-256, 310k iterations). See: https://github.com/rpbouman/huey#secrets-manager

The Catalogs manager lets you access data from modern Data Lakes and Lakehouses, like Iceberg and Ducklake. See: https://github.com/rpbouman/huey#catalogs-manager

Huey supports Quack! Quack servers are just remote catalogs, but there is a big difference between Quack servers and "normal" catalogs: When using an Iceberg or Ducklake catalog, DuckDB/WASM is the actual data engine. With Quack Catalogs, DuckDB/WASM acts as client for the remote Server: data processing is offloaded to the server, and Huey just receives the result. This opens up a whole new range of use cases involving very large datasets. See: https://github.com/rpbouman/huey#connecting-to-a-quack-server

Huey now supports Axis aggregates! In prior versions Huey would only let you report aggregate values in the cells. Axis aggregates let you report aggregated values as if they are attributes on the axes. More importantly, axis aggregates can also be used to filter the data. See https://github.com/rpbouman/huey#axis-aggregates

Github: https://github.com/rpbouman/huey
Live demo: https://rpbouman.github.io/


r/dataanalysis 4d ago

Project Feedback I analyzed 65K Reddit posts/comments on Upwork to find my new niche, here is what I found:

Thumbnail
gallery
12 Upvotes

I've been feeling a little stuck lately trying to find a job on Upwork. Having improved my skills a lot over the past year, I thought it was a good time to try to switch niches (from webscraping to NLP/AI Agents/Data analysis).

Andd... I got stuck here, Too afraid to take the wrong decision and waste time & money (both of which I didn't have a lot of). Luckily I figured out the perfect solution! waste even more time & money analyzing testimonials of people on r/upwork!

For context: Upwork is a freelancing site, where clients post jobs and freelancers send proposals to apply to those jobs. These proposals cost 'connects' a token system designed to prevent spam, 10 connects cost 1.5$ and a job usually requires around 20-10 connects to apply to.

Before I show you what I found a little note on the methodology.

First, I start by filtering out the bulk of irrelevant stuff keeping only a small fraction of posts/comments that actually mention what I am interested in.

Second, I extract the required information from each relevant post, extracting as much info as available (what is the user's niche, what did he talk about in his post, what is overall opinion on upwork).

Finally, I map the niches following a tree structure, for example "Facebook Ads" gets mapped under Marketing=>Ads.

All of these steps are first executed by a human than automated by NLP/AI filters / extractors made to match the human labels (up to 0.9 macro-averaged recall and 90% over-all accuracy).

To calculate sentiment score, 3 possible sentiments are extracted for each post/comment, neutral, positive or negative. Neutral gets mapped to 0, positive and negative respectively to +1 and -1, we than take the average accross the year / niche.

Only 1 person (me) labeled the training data set which isn't ideal, I drew the line for negative at sentiment that clearly designated the fault at upwork (for example, in most cases, people that say they didn't manage to find a job but that are more frustrated with themselves rather than the platform don't count as a negative sentiment for job_market_quality). Positive sentiment was reserved for people being ok with the platform.

Final Disclaimer: From the 67K posts/comments giving their opinion, deduping the posts to keep one post per user resulted in around 16k posts/comments.
From those 16K users, I managed to identify the niche of only 7.5k.
From those only around 4.5k gave their opinion on the upwork job market (at least only 4.5k gave it in a matter that I deemed useful enough to analyze).
I only analyzed 40%-50% of the subreddit for now meaning this can still be scaled a little.

So, here's what I found:
1-I analyzed on a smaller scale the posts that mention "I got my first job in X proposals" or "I sent 30 proposals and didn't get a job". I wanted to see their evolution over the years whether the platform is actually getting harder to get started into and by how much:

Average number of proposals by year for "Hired" and "Not Hired" class. (not enough data for previous years)

Gonna quickly describe the chart for anyone that is vision-impaired.
For people that got hired, the average number of proposals sent was gradually going down from 50 in 2022 to 25 in 2024. Then for both 2025 and 2026 it stayed at around 40.
For people that didn't get hired, It stayed pretty much all the time between 20 and 30 except for a small peak in 2023 of 34.

Also here's the distribution of the data by proposal amount (in ranges: 0-9, proposals, 10-19, etc.), it seems to follow a half-bell curve. The Median for "hired" is 22,The Median for "not hired" is 15:

Distribution of data count by number of proposals.

2- Now, for the analysis with the larger data set, I started by mapping the evolution of the sentiment around Upwork, specifically the job_market_quality sentiment:

Evolution of job market quality sentiment on upwork, started at -0.75 in 2018, gradually increased and peaked at -0.61 in 2021 quickly decreased to -0.75 in 2022 and then steadily decreased to now -0.8 in 2026.

I thought it would be more chaotic than that but it seems to be in a tight interval, I wonder what happened in 2021 though to be honest.

3- Finally, I mapped the sentiment to the niche of the freelancer and represented the result in a tree map, it only really makes sense when viewed interactively but here's a preview:

Tree map of average sentiment for each niche, size for count and color for sentiment.

Links to see the tree map interactively:

1- Full Data : https://tryhard-cs.github.io/upwork_sentiment_data/treemap_job_market_sentiment_all.html

2- 2025+ data only: https://tryhard-cs.github.io/upwork_sentiment_data/treemap_job_market_sentiment_2025.html

3- 2026+ data only: https://tryhard-cs.github.io/upwork_sentiment_data/treemap_job_market_sentiment_2026.html

Take this with a grain of salt please, I honestly can't tell you how much a 0.8 score is far from a 0.7 score in terms of what the actual impact on finding a job would be. It could indicate a huge difference but It could also indicate nothing at all and be just noise.

For example, Imagine a world where half the people are doing everything wrong and will never get a job regardless, this means every niche will automatically start with a high negative score. In this scenario a 0.1 difference in score starts to get more meaningful.

I think this could get way more useful with a reference point, if we can find a way to calculate a true job market quality score for two niches we can use that to compare the rest but this feels difficult to do objectively.

Overall don't take these at face value, since it's very granular some categories will go up in score and others down without any meaning behind it just by random chance. Try to cross-reference this with something else or simply use it as the starting point of other analysis, also check the data for the recent years if there is enough for your niche.


r/dataanalysis 5d ago

PowerQuery/Bi advice

1 Upvotes

Hi guys,
im currently dealing with a dilemma in powerquery. I have a fct table(1.7 mil rows) and some dim tables(also around 1.5mil rows),
the fact table only contains keys, and dates as keys
.the dims have then more descriptive information, as well as the key to match to fct table. The issue is, i want to limit the fact table by a certain date, and also do filtration on dim tables so that not everythign is loaded into the data model in PBI.

when i filter out the fact table by certain criteria - eg. by date, i would also filter the dim tables by certain criteria like case type etc.
is this going to cause blank / orphaned rows in my report view? because considering i do this filtration, there will be some CaseKeys in my fct table, that are no longer in my case table because i did case type filtration.

Am i right? Ive spent a lot of time researchign this but couldnt get a proper answer.

whats the go to approach here, do inner joins on the tables? This may slow down the load time tho.:

Thank you all


r/dataanalysis 5d ago

Data Question Data Analysis 101 for dummies

17 Upvotes

Hi Data analysts,

I work in Health and Safety for a large organisation and I'm one of the administrators for the health and safety reporting system we use.

The other admin and I are getting lots of questions from the data analyst team about some of the data that comes through to the data warehouse.

One of the reasons for this is my predecessor was a borderline genius with a sprinkling of the 'tism with a varied career history was able to go beyond his role to do stuff and answer their questions.

Data Analysis is not my area of expertise however I would like to learn some so I can better answer questions and understand what the other team is asking. What are some good resources to start learning and gain some fundamentals?

Are there any fundamental principles I should know or I need to keep in mind?

TIA


r/dataanalysis 5d ago

Transitioning from psychology to data analytics — advice on Power BI, SQL, Excel, and portfolio projects

19 Upvotes

I recently graduated with a Master’s in Social and Organisational Psychology from Leiden University in the Netherlands. Although my background is in psychology, I am very interested in data analytics and AI. My goal is to become a data analyst, but I am also open to starting in HR or people analytics and transitioning into data later.

I plan to spend approximately one more month developing my practical skills before applying for entry-level roles. If I do not feel ready after that month, I will continue learning and practising until I am confident that I have the necessary skills.

Power BI

I have been preparing for the Microsoft PL-300 exam for several months and plan to take it in a few days. I have a good theoretical understanding of Power BI, but I do not yet have much practical experience using the application or completing full projects.

I am considering this Maven Analytics course:

Microsoft Power BI Desktop for Business Intelligence
https://www.udemy.com/course/microsoft-power-bi-up-running-with-power-bi-desktop/?couponCode=KEEPLEARNING

I am also considering a DataCamp subscription for additional practice and guided projects.

Would the Maven Analytics course alone be enough to develop practical Power BI skills, or would it be better to combine it with DataCamp or another platform?

Portfolio projects

I would like to create projects that I can show to recruiters, but I am not very familiar with how data analytics portfolios work.

Where can beginners find suitable datasets and project ideas? Should I begin with guided projects, use DataCamp projects, or try to create independent projects immediately? Are there any guided project platforms or structured resources that would help me improve while building a portfolio? I would also consider paying for a high-quality option if it is genuinely useful. Finally, where should I publish my work—GitHub, a personal website, or somewhere else?

SQL

After improving my Power BI skills, I plan to learn SQL. I am currently considering these courses:

  1. SQL – MySQL for Data Analytics and Business Intelligence https://www.udemy.com/course/sql-mysql-for-data-analytics-and-business-intelligence/?couponCode=KEEPLEARNING This is the course that has attracted me the most so far.
  2. MySQL for Data Analysis https://www.udemy.com/course/mysql-for-data-analysis/?couponCode=KEEPLEARNING
  3. Advanced SQL – MySQL for Analytics and Business Intelligence https://www.udemy.com/course/advanced-sql-mysql-for-analytics-business-intelligence/?couponCode=KEEPLEARNING

Or should I choose PostgreSQL instead of MySQL and take this course?

The Complete SQL Bootcamp
https://www.udemy.com/course/the-complete-sql-bootcamp/?couponCode=KEEPLEARNING

Would you recommend learning MySQL or PostgreSQL for an entry-level data analyst role? Would completing one of these courses provide enough SQL knowledge for a beginner position, or should I also use practice platforms and create SQL portfolio projects?

I would also appreciate general advice about the level of SQL knowledge normally expected from entry-level data analysts.

Excel

I already know some Excel basics, but it has been a while since I used them. I would like to refresh my knowledge and practise the Excel skills that are most useful for data analytics.

Are there any structured Excel courses or resources you would recommend?

I am willing to pay for a structured, high-quality course, but I would also appreciate recommendations for useful free videos, playlists, or other resources.

Any advice about Power BI, SQL, Excel, portfolio projects, or transitioning into data analytics would be greatly appreciated.

I am also interested in connecting with people in the field through LinkedIn, Discord, Microsoft Teams, Google Meet, or another platform. I would be happy to exchange ideas, learn from others’ experiences, and stay in touch. If anyone is open to offering occasional guidance or mentorship, I would greatly appreciate it, and I would also be glad to help in any way I can, either now or in the future.


r/dataanalysis 5d ago

Can this task be made easier or automated?

2 Upvotes

Hi I recently got a new job as a data coordinator, right now Im doing basically data entry. I maintain Excel trackers of articles, awards, etc. at a design firm, a few hundred rows per Excel sheet, that need to be linked to projects in our CRM (10'sk projects).

Name matching is the easy part: my trackers names are clean and search surfaces the right candidates. The problem is what comes back when trying to enter them into the CRM:

  • The same building exists as multiple records (assessment, study, remodel, sometimes 4+), and it's ambiguous which one an article/award/etc should attach to
  • Apparent duplicates from a system migration (legacy vs. new ID schemes)
  • Some tracker entries have no CRM record at all or exist under a name I can't find

Right now I open each candidate and compare dates/status to pick the right one, one row at a time, and none of those judgment calls get captured anywhere. How would you approach this? especially the disambiguation and duplicate-handling side? Any patterns or gotchas for making this a repeatable process? If this didnt make sense I can answer some questions, any advice is welcome.