The role involves performing initial diagnostics, localizing issues, and owning incidents from detection through resolution. You will collaborate with the team to maintain runbooks and escalate complex issues to L2/L3 engineers.
This isn't classic ticket support. We're looking for someone who becomes the platform owner in the moment of an incident, not just passing the problem along, but understanding what's actually going on, driving communication through to full resolution, and being able to clearly explain what happened and why.
If you've written bug reports that developers could act on immediately without follow-up questions, that skill translates directly here - just applied to production incidents instead of test cases.
Requirements
Experience working with monitoring and logging systems
Ability to read and analyze logs (Grafana, Kibana, Loki)
Basic issue localization (network, DNS, service connectivity)
Basic understanding of Kubernetes: kubectl logs, kubectl describe, kubectl get
Understanding of application configuration: Helm values, ConfigMaps, environment variables
Deep knowledge of our specific platform isn't required going in — we'll train you. Kubernetes/Helm experience is a plus but
Background in QA/tech support
Nice to Have
Experience with Helm
Familiarity with CI/CD pipelines
Understanding of microservices architecture
Soft Skills
Ability to ask precise clarifying questions
Structured problem description when escalating
Independence and ownership
A genuine desire to understand the issue, not just pass it along
Growth-oriented mindset — this is an entry point, not a final destination; our platform evolves fast, and the role has a natural path toward DevOps/development for those who want it
Willingness to work night shifts covering European hours; candidates based near the EST timezone are preferred
Responsibilities
Receiving and processing requests via Telegram, Slack, and email
Initial diagnostics: clarifying the nature of the issue and gathering details from the user
Issue localization: identifying which component or service an incident relates to
End-to-end incident ownership: not just escalation, but owning the incident from detection through resolution, keeping stakeholders updated along the way
Basic infrastructure troubleshooting: logs, service status, configurations
Executing deterministic runbook actions (e.g., telephony service failure identified → restart service → test call → confirm recovery)
Escalating to L2/L3 (DevOps, developers) with prepared context
Creating tickets and bug reports in the tracker
Collaborating with the team to build and maintain runbooks and knowledge base entries for recurring issues
What we offer
The team has built award-winning AI products for tech corporations - devices, voice assistants, products that are actually in the world
Cutting-edge tech stack: Speech Technologies, NLP, Generative AI (LLMs, diffusion models), voice-first agentic architecture with privacy-first and on-premises deployment
High engineering bar and real ownership - the team cares about what actually works in production, not what looks good in a demo, and you'll see the impact of your work directly
Fast career progression - a senior-heavy team and a high volume of real problems means you grow faster than you would anywhere else
Startup pace with enterprise stability - real clients, real revenue, no bureaucracy
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”