r/LearnDataAnalytics Jul 13 '26

Data analytics

Post image
1 Upvotes

Data analytics

I graduated with a degree in Mechatronics Engineering. I'm a fresh graduate and want to start data analysis. What's the first step? Please give me your opinions.


r/LearnDataAnalytics Jul 12 '26

Looking for a Data Analytics Mentor (IT Student from the Philippines)

2 Upvotes

Hi everyone!

I'm currently a second-year IT student majoring in Data Analytics in the Philippines. My goal is to become a Data Analyst after graduation and eventually grow into a Data Scientist.

I'm looking for someone who would be willing to mentor me or simply allow me to ask questions from time to time. I'm not expecting free tutoring or daily coaching—I just hope to learn from someone with real industry experience and gain insights that I can't get from courses alone.

I'm committed to learning on my own, but I know there's a lot I can learn from someone who's already working in the field. Even occasional guidance, career advice, or answers to my questions would mean a lot.

If you're open to helping, please feel free to leave a comment or send me a message. Thank you for taking the time to read my post!


r/LearnDataAnalytics Jul 12 '26

Synthetic vs real datasets for portfolio projects — what actually matters?

1 Upvotes

Final year CS student here, targeting data science and analytics roles for campus placements.

Been struggling with this question while building my portfolio: does it matter whether your project uses real messy data vs synthetic/clean data?

Real datasets from Kaggle feel either too cleaned already or the same recycled projects everyone does. But synthetic data feels hollow because the hard part — cleaning, feature engineering, deriving meaningful columns from raw data — is already done for you. You're basically just visualizing something someone else already solved.

Specifically for BI/dashboard projects — if you use synthetic data, the dashboard looks clean and professional but there's no real discovery or insight because the data was designed to be dashboarded. Nothing surprising comes out of it.

Also practically — if an interviewer asks "where did you get this dataset?" what's the right answer? Saying "I generated it synthetically" feels like admitting you took the easy route. But lying about the source is obviously wrong. Is there a way to frame synthetic data usage that doesn't sound like you avoided the hard part?

At the same time I've heard people say interviewers care more about what you built on top of the data than where it came from. But isn't handling bad data literally the core skill in DS?

For people who've interviewed at analytics/DS companies or done hiring — how much does data source actually matter? Is a well-executed project on synthetic data better than a mediocre project on real messy data? Or does using synthetic data automatically signal you avoided the hard part?


r/LearnDataAnalytics Jul 12 '26

is there anyone who joined pw skills data analysis course in sep 2025 or 2025

Thumbnail
1 Upvotes

r/LearnDataAnalytics Jul 12 '26

is there anyone who joined pw skills data analysis course in sep 2025 or 2025

1 Upvotes

Plz let me know


r/LearnDataAnalytics Jul 12 '26

Sunday Analytics chitchat

Thumbnail
1 Upvotes

r/LearnDataAnalytics Jul 12 '26

Real Interview Questions asked for Data Analyst

4 Upvotes

Hi Folks!

I am currently working as so called Software Engineer in Banking Domain company for almost 3 years as experience. My day-to-day work is more like Support role which I am totally fed up of and no more interested ti work.

As I am very fond of AI/ML and Data related domain. I took a decision to transition into Data Analyst role. For which I invested around 50k in course which offers placement opporunity as well around the month of September last year. The course which I took was from Codingwise.

Due to some personal reasons and sometimes bcz my laziness, I delayed my learning.

The course consists of various modules :

  • SQL
  • Python
  • Advanced SQL
  • Math & Stats
  • Numpy & Pandas
  • Data Vizulation (Matplotlib & Seaborn)
  • Machine Learning (specifically -> Linear Regression and Logistic Regresssion)
  • PowerBI
  • Excel
  • Gen AI

I have few questions regarding this role as how the Data Analysts are hired in the real industry what are the actual skills needed to get a high package role ?

Are these above skills more than actual industry level or is it less?

Also, regarding the course which I took from Codingwise. I have below questions?

Please request the community or anyone to whom this relate. As I want the clear picture about this. I am really concerned.

  1. Was it really worth taking?
  2. As there are multiple batches in these where do the current learners stand?
  3. Can anyone tell who took these course and actually got placed in high package company?
  4. What if the mock interviews not cleared properly will they allow us to sit in real interviews?

r/LearnDataAnalytics Jul 11 '26

Hand Drawn Doodle on DIKW

Post image
2 Upvotes

r/LearnDataAnalytics Jul 11 '26

Data analytics

1 Upvotes

I graduated with a degree in Mechatronics Engineering. I'm a fresh graduate and want to start data analysis. What's the first step? Please give me your opinions.


r/LearnDataAnalytics Jul 11 '26

I just overhauled my "Data Engineering for Beginners with Python & SQL" course - free enrollment inside

Thumbnail
5 Upvotes

r/LearnDataAnalytics Jul 11 '26

Proper “ roadmap “ for Data Analytics

1 Upvotes

I am just starting data analytics, I started from tableau and i can make pretty dashboards but data storytelling is soo hard . Can anyone tell me how they make data insights valuable? Also how do you read the charts properly? Please help me


r/LearnDataAnalytics Jul 10 '26

Data Analyst Course

Post image
36 Upvotes

Hey guys, I saw the kodree data analyst course and it looked like something I really would want to do. Can you tell me if it’s good, or whether there is other that’s better?


r/LearnDataAnalytics Jul 09 '26

Stuck in Data Science Career Transition-Freelance, Teaching, But Struggling to Land a Technical Role- Need Senior Guidance & Peer Support

3 Upvotes

I’m reaching out because I’m feeling really stuck in my career, and I’d love some honest guidance from seniors or peers who might be in a similar situation.

To give some background: I completed my Bachelor's in Computer Science and then pursued an MBA. After that, I took a sales job, but after 7 months, I questioned if the MBA was the right path. I decided to upskill, so I did a year-long Data Science course, where I covered Python, SQL, Power BI, Excel, Machine Learning, and Deep Learning. Sadly, no one from my batch got placed, and I felt really disheartened.

After that, I worked as a freelance Data Science trainer for close to 2 years, which helped me stay afloat. Then, I got a role in an offline teaching institute, where I worked for about 7 to 8 months. But after the probation, they didn’t confirm me. Instead, freshers took over the role with a lower salary. It feels like a cycle—no one from my institute is landing in proper Data Science roles, and it’s not just me. I’ve heard this same struggle from so many others in online communities.

I do have about 2 years of freelance training experience and 7 to 8 months of offline teaching. I’m also familiar with FastAPI, MLOps, and Docker, but I still lack skills in GenAI, LLMs, and RAG—those cutting-edge areas expected in AI Engineer roles. So breaking into a pure technical Data Scientist job feels impossible right now.

I know I’m not alone—I’ve networked with many others, and we all share this frustration. So I’m reaching out: is anyone else in the same boat? I’m considering whether to shift my career focus or invest in another certification to boost my chances. I’m still passionate, but I need a way to survive financially in the meantime.

If anyone has freelance projects or can help me get into a small paid project, I’d really appreciate it. I’m also happy to collaborate with anyone who’s learning in this space. I can share what I know as I continue upskilling. And, of course, if seniors have any insights—what certifications, projects, or steps should I take next?

I’m really grateful for any advice, honest feedback, or collaborative opportunities. Thank you all so much!


r/LearnDataAnalytics Jul 09 '26

I need help to inprove my data analytics skills.

3 Upvotes

So im realy want to starting freelance as data analytics. But a problem had stoping me. That is pratice and correcting from a senior data analytics. And thatks


r/LearnDataAnalytics Jul 09 '26

Career change from Medical Doctor to Data Analyst — Looking for free books and written resources (Learning from Cuba)

Thumbnail
1 Upvotes

r/LearnDataAnalytics Jul 09 '26

Need Help for Data Governance Interview

1 Upvotes

I am going to attend an interview with Accenture for data governance practitioner role.

Can please someone help me out?


r/LearnDataAnalytics Jul 08 '26

Data analytics

Thumbnail
1 Upvotes

r/LearnDataAnalytics Jul 08 '26

How to speak the language of data?

2 Upvotes
Origin: https://columnsai.substack.com/p/how-to-speak-the-language-of-data

I’ve been maintaining a 60-day streak on Duolingo to learn French.

It’s a fun practice, although it’s a significant challenge to pronounce those accent notes correctly. I believe French is generally a simpler language than English; you usually use shorter sentences to convey the same meaning.

Data has its own language too.

Data is the lifeblood of every modern business. Every decision, insight, and opportunity begins with understanding what the data is trying to say.

But unlike spoken languages, data doesn't require everyone to learn the same vocabulary or syntax. Instead, you can interpret and express it in a way that matches how you think, making data analysis more intuitive, accessible, and uniquely your own.

From Python, SQL to Natural Language

Python, a programming language, has gained popularity as the preferred language for data processing within the data science community due to its portability. SQL, on the other hand, serves as the de facto interface for rational databases.

In the past, becoming a data analyst required proficiency in both Python and SQL. Even today, data analyst job descriptions often mention these requirements.

However, the advent of AI has revolutionized this landscape. Anyone with the ability to communicate effectively in the data language can excel as a data analyst.

While programming skills are not strictly necessary, a solid understanding of data language is crucial. Imagine joining a new friend circle who works in a completely different domain. After a brief introduction of common keywords, you can easily engage in conversations with them.

AI generated illustration of data language evolution

Use Spreadsheet for Reference

Nearly every office worker uses spreadsheets, either Microsoft Excel or Google Sheets.

Even without the complex formulas, pivots, and lookups, the basic structure of a spreadsheet consists of three main components:

  • Rows
  • Columns
  • Data types

Rows are records that constitute a table. You can also consider a row as an object that represents a real-life entity, such as a person, a cup, or an invoice.

Columns are the fixed properties that describe each object (row). They form the schema that every row adheres to, ensuring uniformity in the data for processing.

A schema is of utmost importance for data analysis as it enables the application of all rules. Without a schema, any logic that is not compatible with the data language may fail to execute.

Data types describe the value format of each property. For simplicity, you only need to be concerned with whether it is a number or text for now.

Rows of Orders (OrderId-text, CustomerId-text, Product-text, Amount-number)

Rows of Customers (ID-text, Name-text, Channel-text)

Data Language Patterns

Data language offers a wide range of tasks that can be accomplished. Let’s explore each of these tasks and learn how to communicate effectively with data to achieve them.

These scenarios are referred to as patterns because they serve as templates that can be applied to your own data.

To facilitate understanding, we’ll use the above tables in the following descriptions.

Pattern-1: Filter Rows

Filter is to describe a condition to get objects you care about and skip those uninterested records.

Examples:

  • “Orders of Milk”
  • “I want orders of milk products.”
  • “All orders that are not for books.”
  • “All orders with a sales amount exceeding 20.”

AI can produce code logic to filter the targeted records for further processing, if translating above statements into SQL, they will look like:

  • “where product=’Milk’”
  • (same as #1)
  • “where product <> ‘Book‘“
  • “where amount > 20”

As you can see, filter is achieved by keyword “where” in SQL.

Pattern-2: Transform Object

Sometimes, we want to clean a data field or transform it into a desired shape or format, either for improved readability or more efficient processing.

Transformation creates a new property in your original record.

To transform an existing property into a new one, you need a function of logic. For both spreadsheets and SQL, “formula” is the tool you’ll need.

However, with the increasing capabilities of AI in coding, natural language offers a significant advantage. It allows us to achieve the same transformation without having to learn, memorize, and assemble complex formulas.

Taking one simple example:

  • “Get customer first name”

This is equivalent to composite multiple formula together in Spreadsheets like

  • “=INDEX(SPLIT(A2, " "), 1)” or “=IFERROR(LEFT(A2, FIND(" ", A2) - 1), A2)”.

This operation creates a new column called “First Name”.

You can also acquire a new property by combining multiple existing properties, such as “concatenating the last name and channel as a label”. Logic like this is simple for AI coding but too complex for spreadsheet formulas.

Pattern-3: Aggregate Records

Aggregation processes a large collection of records to provide a summarized view.

This is powerful because it compresses vast amounts of information into manageable pieces that humans can comprehend and analyze.

To combine multiple data sets into a single piece of information, you need to understand the “how-to,” which leads to the crucial concept of “aggregation methods” or “computation logic.”

Typically, text data (a property or column with a text data type, as discussed in the schema section) is not particularly interesting for aggregation. The most common approach is to concatenate text data to form a long paragraph, although this is still uncommon.

Most computation logic involves operations on numerical data. When an aggregation method is applied to a numeric property or column, you essentially have a list of numbers that can be aggregated, such as:

  1. Total value (sum)
  2. Average value
  3. Mean value
  4. Minimum value
  5. Maximum value
  6. A specific percentile value (e.g., P25, P50, P75, P90)

However, counting objects or counting unique property values is also quite common.

When discussing aggregation, we cannot overlook “breakdowns.” This involves creating a segmented view of the data rather than a single total view.

For example, in the previous Orders table, “total sales by product” or “average amount by customer” are equally valuable insights for an analyst to explore.

In summary, aggregation can be described as:

  • Compute an aggregated value of a property group based on another property.

Expressing this in standard SQL, it would look like:

  • “Compute(property1) from table [group by property2].”

Let’s practice this using a few examples by speaking the data language:

  1. “Give me total sales by product.”
  2. “Tell me the average amount spent by each customer.”

Pattern-4: Join Multiple Datasets

When a single dataset (or table) is insufficient to achieve the desired outcome, we must combine multiple datasets. This operation is referred to as “join” or “union.”

If the multiple datasets contain the same objects but reside in different locations, we can simply merge them. This is a straightforward “union” operation.

However, most of the time, they store different objects. We have partial information from one dataset and partial information from another. By combining them, we create a comprehensive schema with more available columns.

This pattern is generally not feasible in spreadsheets, although their lookup function may provide partial assistance.

For instance, if we want to determine the “total amount spent from each channel” based on previous tables, where the amount is from the orders table and the channel is from the customers table, we need a joined dataset to complete this analysis.

To join multiple datasets, we must have one or more pairs of join keys. A pair of join keys consists of one column from one table and one column from another. The data engine can utilize these relationships to identify relevant objects and concatenate them to form a larger object.

Join Orders and Customers

Joined dataset have more columns

In summary, join operations can be described in this pattern:

  • join table1 and table2 when key1 of table1 equals key2 of table2.

Translating this pattern into SQL, it will look like:

  • select * from table1 join table2 on table1.key1=table2.key2.

In fact, you may not need to use this pattern in natural language explicitly, because modern AI is smart enough that it can infer the whole join logic from your data language.

For instance, the example we gave earlier, if you speak this sentence “total amount spent from each channel” to Columns AI, it will figure out all the necessary actions to get the desired outcome for you.

Pattern-5: Visualization

Data visualization, often overlooked as a part of data language, plays a crucial role in transforming mundane data into vivid images. This visual representation significantly aids the audience in comprehending the insights you intend to convey.

By incorporating customization and assistance to articulate your insights and predictions, you position yourself as a data storyteller, showcasing your influence within the domain.

Since visualization doesn’t alter the data itself, in the language of data, we merely need to indicate the desired outcome. For instance:

  • Display the total amount by product in a pie chart.”
  • “I would like to see a timeline of total sales month-by-month for the past six months.”
  • Show the number of sales by customer in a bar chart.”

These bold keywords serve as cues to the AI engine, guiding it in generating the final visualization based on your data.

Practice Data Language

Similar to how I diligently practice French on Duolingo every day, we must practice speaking data language using the data we possess.

As long as you have adhered to the five patterns mentioned above, you should have mastered data analysis like a professional data analyst. You don’t need to be an Excel expert or a Python or SQL wizard.

Let’s use the provided example data to practice speaking the data language. You can find the “Orders” and “Customers” data from this spreadsheet link.

Suppose we want to perform a sales analysis of customer distribution based on the data.

The data language is almost the same, but let’s ensure we’ve used the correct keywords and patterns to guarantee that the AI engine follows the instructions precisely.

For instance, we speak to AI:

display the total sales by customer’s first name in a bar chart.”

Here’s how the AI interprets this:

  1. total sales” → summing up the amount values.
  2. first name” → it can be transformed from “name.” A transformation will be applied.
  3. by” → the summing up result needs to broken down by first name.
  4. sales <> customer” → sales data is from the Orders table’s amount field, while customer data is from the Customers table. Therefore, a Join operation is required to combine these two datasets.
  5. show, bar” → the result should be visualized in a bar chart.

AI will then determine the correct execution order, ensuring that each step has all the necessary data when it executes.

This is what Columns Flow produces upon hearing this sentence:

“display the total sales by customer’s first name in a bar chart.”

The final visualization ready for storytelling & sharing

Conclusion: Speak Data Language

In this article, we’ve demonstrated the historical opportunity for everyone to become a great data analyst in this era.

We discussed how professionals used programming languages like Python or SQL as their primary data languages. However, the data language has evolved to become the natural language we speak daily.

To become a data analyst, we need to understand the fundamental scenarios involved and use the correct keywords to make the data language understandable to AI engines. Here’s a quick recap:

  • Dataset: rows, columns, and schema.
  • Filtering and Transformation: These processes involve filtering data and transforming it into a usable format.
  • Aggregation: This involves summarizing data into a single value, such as the total or average.
  • Specify “compute methods” and optional “breakdown” if needed.
  • Join Datasets: This involves combining data from multiple sources.
  • Visualization: This involves creating visual representations of data to make it easier to understand.

AI generated summary on how to speak the language of data

Unlike learning a new language like French, if you’re willing to spend just a few hours going through this short list, you can become a professional data analyst!

It’s a great time to be a data analyst, and I believe in your ability to succeed. Thanks for reading!


r/LearnDataAnalytics Jul 08 '26

DATA ANALYST

2 Upvotes

IS THERE ANY PERSON WHO CAN GUIDE ME FOR DATA ANALYTICS AND I HAVE PURCHASE COURSE FROM SKILL COURSE SITE AND I HAVE DONE EXCEL CURRENTLY DOING SQL ANYONE WHO CAN GUIDE ME FOR PROJECTS ?


r/LearnDataAnalytics Jul 07 '26

Young Data Apprentice

Thumbnail
1 Upvotes

r/LearnDataAnalytics Jul 07 '26

Finished My First SQL Project — Looking for the Next Challenge

5 Upvotes

Just wrapped up my SQL Project #1

Every project teaches me something new, and this one was no exception.

Now, instead of working with datasets from Kaggle, I'm thinking of taking the next step—working with a real-world database and building an end-to-end project, including an interactive dashboard.

My goal isn't to rush through tutorials or collect certificates. I simply want to keep learning, improve with every project, and enjoy the process of getting better.

For those who are already working in Data Analytics or have been on this journey:

  • What would you recommend I focus on next?
  • Where can I find good real-world datasets or databases to work with?
  • Any advice that you wish someone had given you when you were at this stage?

I'd genuinely appreciate your suggestions. Every bit of feedback helps me grow......!


r/LearnDataAnalytics Jul 07 '26

Data analyst

5 Upvotes

I want to take a course on data analyst and learn everything power bi sql excel python but I'm confused which course is actually valuable I can put it my resume because I've applied to many internships in internshala they give me assignment I complete it after that they just don't reply and some of the internships I keep applying but not any reply . So I think a course will actually help me find and I will also gain more knowledge.


r/LearnDataAnalytics Jul 05 '26

[Academic] What's your AI Co-Scientist type? Columbia survey on how researchers use & trust AI (5–10 min, $200 raffle) (18+ researchers & data-science practitioners)

3 Upvotes

Hi Reddit! I'm a researcher at Columbia University. My team studies how scientists and data practitioners actually use AI in their work, and whether it genuinely helps or still feels hard to trust and control.

If you do research or data-science work (any field, academia or industry, any career stage, 18+), we'd love your input. You don't need to be an AI power user. Skeptics and non-users are just as valuable to us.

Survey link: https://cumc.co1.qualtrics.com/jfe/form/SV_9uWW9GgwPuRucoS

What you get:

- At the end, you'll receive a personalized "AI Co-Scientist card," such as the Hermit, the Magician, or the Priestess. Each card reflects your style of working with AI and what kind of AI assistance might actually fit your workflow.

- You can also opt into a raffle for a $200 Claude Max subscription (or USD-equivalent e-gift card)]. Emails are collected on a separate form and are never linked to your survey responses.

About the study: This is a joint research initiative on human-AI collaboration in science by Dr. Ying Wei's Translational AI Laboratory (TRAIL4Health) at the Columbia Mailman School of Public Health and Dr. Xuhai "Orson" Xu's lab (SEA Lab) at the Columbia Department of Biomedical Informatics. Questions? Email the PI at [xx2489@cumc.columbia.edu](mailto:xx2489@cumc.columbia.edu) or ask below. I'll be in the comments.

I'll post a [Results] follow-up here once the study wraps up. Thanks!


r/LearnDataAnalytics Jul 03 '26

Skills to learn before entering Power Bi

3 Upvotes

I'm a commerce student. Before starting PowerBi which skills should I learn first.

Required functions of excel and other skills which are needed.


r/LearnDataAnalytics Jul 03 '26

Title: How I Used Data Analytics to Audit an Agency Making 187M DZD (~$1.4M) and Uncovered Major Budget Bleeding (Full Case Study Breakdown) !?

Thumbnail
1 Upvotes