Match your resume skills with our AI powered skill match!
You will be responsible for redesigning and implementing the observability stack to ensure visibility into system performance and reliability at scale. This involves hands-on systems engineering, from data collection in code to processing, routing, and visualization.
Whatnot is the largest live shopping platform in North America and Europe to buy, sell, and discover the things you love. Whether it's trading cards, fashion, electronics, or live plants, our sellers are building real businesses across hundreds of categories. We're building live commerce at a scale that's never been done in the West, and there's no playbook to copy. The people here are shaping how an entirely new industry develops.
As a remote co-located team, we're inspired by our values and anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia. We move fast, stay close to our users, and focus on the work that drives the most impact.
We're one of the fastest growing marketplaces and were recently named the #1 Best Startup Employer in America by Forbes. Check out the latest Whatnot updates on our news and engineering blogs and join us as we enable anyone to turn their passion into a business and bring people together through commerce.
The Infrastructure Reliability Engineering team is looking for seasoned Software Engineers who will be responsible for rethinking and redesigning how we do Observability at Whatnot. As our scale, traffic, and complexity continue to grow; yesterday’s tools, vendors and platforms are becoming obsolete and impractical. Your job is to ensure that we have visibility into the state, performance, reliability and user’s experiences of our software stack, and that this visibility remains 10x and 100x our current scale.
This is hands-on, software-first, systems engineering work. You will work closely with the Core Infrastructure, Platform and Developer Tools teams to redesign and implement every step of observability, starting from collecting data in the code, through normalization, routing and processing, to querying and visualisation. We’re at the scale where it’s no longer reasonable to use simple, off-the-shelf solutions. You will have to use a combination of analytics engines and vendors to make sure we are able to predict incidents, and should they occur, have a swift path to resolution. At Whatnot, we are utilizing AI agents heavily during troubleshooting – our observability stack should take leverage of that.
You will be working on creating a platform that:
Can identify leading indicators for issues before they become a problem
Understands that not all data is equal and can work both with high-volume, low-signal data and low-volume, high-signal data
Uses industry good practices – based on open standards (Otel, Semconv)
Utilizes any set of open-source software or vendors to achieve its goals
Interfaces with infrastructure to provide better signals for infrastructure management than CPU utilization
Correlates events from infrastructure, CI/CD pipeline, experimentations or load tests with logs and time-series data to quickly identify “what has changed?”
Provides good experience for a wide range of users: application engineers, incident responders, leadership
Aids in driving reliability and experience for our customers, staying reliable and effective itself
This is a highly visible role. The Reliability team provides foundational systems and frameworks that allow Whatnot to scale rapidly while remaining stable and trustworthy for buyers and sellers.
People who do well at Whatnot tend to be comfortable figuring things out as they go, biased toward action, and genuinely curious about what they're building. They care more about outcomes than credit and stay close to the product and the people using it.
7+ years of experience designing and building large-scale distributed systems. Experience in Python, Elixir, or Go is preferred, but strong engineers from other backend stacks who are eager to learn are welcome.
You identify as a software engineer first. You want to build systems and write code, not just configure infrastructure or respond to pages.
Have a strong understanding of observability principles, knowing different types of metrics, logs, events and traces; and where they are applicable
Strong fundamentals in designing, building, and operating shared production services and frameworks.
Experience with one or more of the following:
OpenTelemetry
OLAP databases
Creating developer-facing tools, libraries, frameworks
High-traffic, real-time, or event-driven systems
Comfortable in cloud-native environments such as AWS or GCP with Kubernetes and infrastructure as code.
Strong collaborator with clear written and verbal communication skills.
Whatnot is proud to be an Equal Opportunity Employer. We value diversity, and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, parental status, disability status, or any other status protected by local law. We believe that our work is better and our company culture is improved when we encourage, support, and respect the different skills and experiences represented within our workforce.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”