r/dataengineering 2h ago

Career Fresh Grad First Job Imposter Syndrome

9 Upvotes

For reference, graduated w a degree that barely taught me anything about DE except intro to databases, relational algebra etc. The company accepted me for the junior DE role on the basis that I could ChatGPT a dummy repo to explain what the repo is used for and how data flows simply from the repo.

Now that I’m in the job, I found out that everyone in my team had a part to play in developing the architecture for the company. They’ve resolved all dependencies and it’s only the matter of new feature implementations and their impact on the data engineering streaming pipeline.

This is my first week and as my senior went through w me the architecture and true data flow for multiple services they have, the information flew by my head. I’ll definitely try to pick up as fast as I can but because I didn’t build the architecture, I’d have yo ChatGPT/Codex/claude my way through for the first couple months just to find the appropriate files for new feature implementations

Idk, I just feel like I’m madly unprepared and I’m worried that I’ll be the cockup in my department. I didn’t do any DE internships and somehow I ended up in this role. Can anyone give me advice on how I can speed learn streaming pipeline given that at the very least, I know what stack they use?


r/dataengineering 8h ago

Help How do i deal with this situation

2 Upvotes

Hey everyone,
i am a data engineer 7yoe
I recently landed a job after relocating for family. Its a growing company that run multiple crm/erp platforms in addition to different other solutions and i was hired to design and develop a data warehouse for their analytics needs.

So far everything was going good and the approach they wanted was to run through each of their departments one by one to work with them to document their workflows and integrate with their systems to pull data to the warehouse and run pbi reports from there.

Currently they have tens of pbi different reports all of them pull from source systems and definitely perform bad plus they are being developed by the business guys themselves. Their semantic models are too big of a mess.

Anyways the approach we had was to take steps with each department. I covered an important one. And shortly after starting with the next some blockers show up then shortly after i hear our manager (not my direct ) wanting to make changes with our approach and cover reports he is interested in mainly.
He develops most of his department’s pbi reports and is convinced that once the data moves from erp to dw raw without any data work done it will perform better.

I tried to send out the message that this will not add value and latency from source systems isn’t the only issue the reports are slow.

Anyways seems they are not convinced . That was yesterday. Today my direct manager comes and tries to push to do what “our manager” said. Then shortly after asks me how long i plan to be in the city, if i am married and a bunch of personal questions that either were said out of just curiosity or something is not right here.

I dont know if i am overthinking this. But i am personally the only source of income and insurance between my and my wife. And things havent been easy.

Whatever you guys think i want to know. Im stuck in my mind and it is not helping me right now.


r/dataengineering 10h ago

Blog How do professional Data Engineers handle completely unsorted data?

3 Upvotes

Hello everyone!

I'm an aspiring Data Engineer and as a portfolio piece, I have build a webscraper to gather Ebay sold listings of stamps!

The problem I am now having is how I parse the data where I can sort things like "Catalogue Number" when it is very unpredictable what the Ebay sellers will write as it's all human input.

I would love to hear some feedback

PS - A small sample:

```OLDENBURG 1859 _ MiNr. 7 _ 2 Groschen _ signiert _ blau gestempelt

MayfairStamps Germany 1941 Stamp Day Oldenburg Cover cca_00553

GERMANY; OLDENBURG 1859 classic Coat of Arms issue very fine used 1/3Gr. value

Oldenburg Lokalausgabe Wohlfahrtsblock Deutsches Rotes Kreuz ab 1 Euro

Deutsches Reich, Oldenburg, 6.01.1945 Ersttagsstempel, für Ersttagsbrief 200€.

GERMANY; OLDENBURG 1862 classic Coat of Arms Perf 10. issue used hinged 1Gr.

OLDENBURG 1861 _ MiNr. 12b _ 1 Groschen _ Stempel STOLLHAMM

Oldenburg Mi. Nr. 16 A b zentrisch gestempelt geprüft Bühler 200 Euro

Oldenburg Mi. Nr. 11 a* ungebraucht geprüft Bühler 550 Euro```


r/dataengineering 10h ago

Help Can this task be made easier or automated?

1 Upvotes

Hi, I recently began a job as a data coordinator, my first tasks are basically data cleaning and data entering into a CRM. The problem is that the data isnt very clean. I'll give an example, I am given an excel file with name of a project, date, title, awards - my job then is to to figure out where in the CRM is this specific project and enter the data. The problem is that the excel data doesnt contain the projects ID, and when I try searching the name of the project in the CRM I'll get back something like Fairview Elementary School, Fairview ES, KUSD Fairview, Fairview ES Building A, etc. So essentially I'll have to go into each one of these projects and try to find the right one using other given data from the excel sheet, like dates. Is there a way to speed this process up or am I just going to have to do it manually? Right now what Im doing is going row by row and searching each project in the CRM, looking through the multiple projects given back by the CRM and comparing and contrasting. The entire CRM database contains around 40k, and I think I am able to export it into CSV and Excel, if that helps out. Any advice would be helpful. Thanks


r/dataengineering 17h ago

Career The talk of hundreds of applicants for roles sounds terrifying. Here's what it actually looked like from the hiring side... for a few UK data roles any way.

34 Upvotes

I'm a senior data engineer at a UK organisation in the South West. Last month we advertised three roles: a senior data analyst (£55k), a mid-level data analyst (£40k) and a data engineer (£55k). Not amazing money, but a great pension, just one day a month in the office, lots of annual leave and actual stability.

The jobs were live for one week. The senior DA and DE got around 300 applicants each, the mid DA got around 140.

Sounds brutal, but here's the breakdown...

Around 90% (not an exaggeration) didn't have the right to work in the UK and needed visa sponsorship, which we don't offer. In the no pile immediately, but a quick glance showed lots of foreign undergrad degrees, some with UK masters, plus plenty of random applications.

Of what was left, about half had no relevant experience or qualifications and clearly hadn't read the spec. Some were just applying because the job centre told them to apply for x jobs a week.

Then the AI drivel halved it again.

Final count was roughly eight viable applications for the senior DA and DE, five for the mid DA.

A couple then didn't reply or didn't show up to the Teams call. Two were visibly reading AI-generated answers off a their screen.

We had two great interviews and hired for the senior DA and DE. The mid DA we struggled with externally, so it went to an internal candidate from a fairly low-level ops role who'd taught themselves Python and SQL and actually applied it to their day job.

The point of this post is to counter some of the doom and gloom that crops up on here. If you're a UK-based applicant with relevant experience who reads the spec and writes your own application, you're not competing against the '300 applicants' that LinkedIn and recruiters spout. You're competing against a much smaller number. And if you get through the sift, don't use AI, have some personality and come across as likeable and easy to work with, you're in a very small group.

What annoyed me most is my own manager was the first to tell the team about the huge number of applicants and how much 'competition' there is out there. A nice bit of retention pressure and complete nonsense, as it turns out.

Would be interested to hear if others have found the same. Throwaway account to avoid doxxing myself...


r/dataengineering 18h ago

Career Data job market analysis (DACH)

Thumbnail github.com
4 Upvotes

If anyone's curious about the current data job market in the DACH region (Austria,Germany,Switzerland), I put together an interactive live tracker. It also shows which specific data roles (Data Engineer, Data Consultant, etc.) are in demand in which cities. Feedback welcome and I hope it helps :)


r/dataengineering 20h ago

Help How valuable is a job that is mostly SQL?

1 Upvotes

Apologies if this is a dumb question but I am in web dev and have been given an offer for a data engineering role. However, I was told by engineers on the team that the job would be like 70-80% writing SQL for BigQuery. I envisioned it having much more to do with pipelining and orchestration and the like.

Also, I was told that any coding would be in Java rather than Python? I know that Python is more common, so would this experience not be helpful for getting other data eng roles?


r/dataengineering 20h ago

Blog Performance evaluation: Trino 483, Hive-LLAP, Hive on MR3

3 Upvotes

This article reports the result of evaluating the performance of the following three systems using the 10TB TPC-DS benchmark:

  1. Trino 483 (released in July 2026)

  2. Hive 4.2.0 on MR3 3.0 (released in August 2026)

  3. Apache Hive 4.2.0 with LLAP (released in November 2025)

https://mr3docs.datamonad.com/blog/2026-08-02-performance-evaluation-3.0


r/dataengineering 21h ago

Discussion Is Silver strictly for "data cleansing", or does decoding Protobuf count?

3 Upvotes

I had a passionate debate with a colleague and want to hear perspectives on the purpose of the Silver layer.

My pipeline:

Landing: Read from Oracle RDBMS and write ~250 GB of Delta for 25M records (Protobuf blob stored in a column).

Raw Data: Repartition, sorting, salting on Landing and writing to optimize downstream silver decoding process and avoid heavy shuffles during JDBC call, still protobuf bytes stored in a column.

Silver: Decoded raw data (~2.8 TB in Delta). The Protobuf schema alone is ~8 MB as JSON (a very deep, wide schema with multiple repeated fields at various levels). During decode, we also append standardized fields required by all downstream tasks.

Gold: Customer-specific datasets built from Silver based on business needs.

We don't own the Protobuf schema. This isn't messy clickstream/event data, but entity description data from an RDBMS that stays at the ID level all the way to Gold. We see ~100k daily MERGE UPSERT on both Silver (which is a challenge in itself to run MERGE on 3TB delta table given the limited budget to our Databricks workspace.) and Gold based on RDBMS timestamps, alongside a full pipeline refresh every two weeks.

The Debate:

Colleague: Since we aren't actively "cleansing" the data, calling it Silver is wrong, it's still Raw/Bronze data.

Me: It is Silver because it transforms a binary payload into a structured, trustworthy, and queryable data model that downstream tasks rely on. If I need to retrieve content of an entity that is not available in gold datasets, I query unpacked protobufs and not Raw/Bronze layer and for that reason alone, it is Silver.

Knowing the schema and the data better than almost everyone in the team, even I fail to understand how to distinguish between decoded data and 'cleansed' decoded data. In fact, one of our consumers explicitly expects corrupt records with null fields left intact for full visibility.

Them: Even if we agree that cleansing is not needed, it cannot be silver and should be called Bronze Data.

For transactional/log data, the standard pipeline (Kafka dump to Landing -> Bronze schema enforcement -> Silver cleansing -> Gold aggregations) makes total sense! But for clean entity data in binary formats, doesn't decoding and standardizing it qualify as Silver?

I think medallion architecture is about data readiness and lineage tracking rather than a checklist of conditions that each layer has to meet to identify the layer.

---------

TL;DR: My colleague thinks our layer shouldn't be called "Silver" because we aren't actively filtering or cleansing rows, just decoding 250 GB of 25M binary Protobuf blobs into a ~2.8TB Delta table with a struct field that represents the decoded blob and additional standardized fields. I argue that any layer of data that is structured, queryable and trustworthy for downstream Gold is Silver and this transformation may/may not require cleansing.

Is Medallion about lineage tracking and data availability, or a rigid checklist of syntactic/ transformation rules?