Hi! I'm Vijay

Software, Data & AI Engineer

I build petabyte-scale data platforms, real-time streaming systems, and AI-powered applications. Currently leading the next-generation data platform at TRM Labs.

Vijay Shekhawat

About Me

I'm a software engineer who loves building data and AI systems. But it all started with a video game.

When I was a kid, my computer crashed in the middle of a game. I was so desperate to get back to playing that I spent hours reading articles online, figured out how to reboot and reinstall the entire operating system, all just to play that same video game. But somewhere in that process, I fell in love with computers themselves. How they work, why they break, and what you can make them do.

That curiosity never went away. It led me to study Computer Science at Shiv Nadar University, and from there into a career building software at scale. I've worked on data systems at LinkedIn and Expedia, and now I'm at TRM Labs, where I lead the engineering of a next-generation data platform as a Staff Software Engineer.

Over the years, I've found my sweet spot in data engineering: designing real-time streaming systems, building lakehouse architectures, and solving the kinds of problems that come with operating at petabyte scale. More recently, I've been exploring the world of AI and finding new ways it intersects with the large-scale data systems I build every day.

I also enjoy sharing what I learn. Whether it's speaking at conferences or writing technical deep-dives, I find that explaining ideas to others is the best way to truly understand them.

For the full career details, check out my LinkedIn.

Conference Talks

I enjoy sharing what I've learned from building large-scale data systems. Here are some of my recent talks.

Beyond the Code

Technical Writing

I write about the architecture and engineering decisions behind large-scale data systems.

Replacing the Engine Mid-flight: Upgrading StarRocks Across 58 Releases With Zero Downtime

How we built a closed-loop AI agent that turns production traffic into automated database optimization, upgrading StarRocks across 58 releases with zero customer impact.

Read on TRM Labs Blog →

Architecting Real-Time Blockchain Intelligence at TRM Labs

How we built an in-house real-time data processing pipeline to reliably process petabytes of blockchain data, comparing Spark Structured Streaming, Apache Flink, Beam, and Kafka Streams.

Read on Medium →

From BigQuery to Lakehouse: How We Built a Petabyte-Scale Data Analytics Platform

The journey of migrating TRM Labs' analytics infrastructure from BigQuery to a modern lakehouse architecture.

Read on TRM Labs Blog →

Apache Iceberg Is Just the Tip of the Iceberg, Part 1

A deep dive into Apache Iceberg's architecture and why it's reshaping how we think about data lakehouse table formats.

Read on Medium →

Mastering Stream Joins in Real-Time Data Processing

A comprehensive guide to stream joins: how they work, when to use each type, and patterns for instantaneous data enrichment.

Read on Medium →

Handling High Throughput Real-Time Pipeline Writes to Databases

Challenges and strategies for writing high-throughput streaming data to databases without losing reliability or performance.

Read on Medium →

Kubernetes Through the Lens of Data Engineering

A data platform engineer's journey with Kubernetes, exploring its capabilities across various data engineering projects.

Read on Medium →

Let's Connect

Whether you want to discuss data engineering, conference speaking, or just say hello, I'd love to hear from you.