SRE — Incidents & Monitoring

 Posted 7 days ago
  
 Brazil
  
5-10 years experience
Apply Now

Please mention DailyRemote when applying

AI Summary

Responsible for managing high-severity production incidents, conducting diagnostics, and performing post-mortems. The role involves designing and maintaining the observability stack to reduce MTTR and increase service availability.

Why work with us? 

 

We are a fast-growing company that is revolutionizing the world of SaaS platform and data in Latin America! 

CIAL Dun & Bradstreet is the leading provider of business decisioning solutions and commercial data across Latin America and the Caribbean. Our solutions are designed to transform how businesses manage risk and make critical decisions about the companies they rely on. 

It’s our people, not technology, that makes what we do possible. It’s our people, not data, that turns information into insights. And it’s our people, not algorithms, that help our clients make better informed decisions. We are innovative, agile, and inspired by SaaS solutions and data – and we are looking for people that share those values to join our mission to build stronger businesses! 

 

About Us 

Dunsguide, by CIAL Dun & Bradstreet, is transforming how Latin American businesses discover and evaluate potential partners. As the region's leading platform for customer prospecting and supplier discovery, we help companies make smarter, data-driven decisions about their business relationships. We're growing rapidly and expanding our suite of solutions that combine deep business intelligence with modern, intuitive tools. 

Sobre a posição

Atuar na linha de frente de confiabilidade da CIAL: resposta a incidentes de produção de classificação mais crítica (highest severity) e monitoramento contínuo de aplicações, serviços e bancos de dados. É uma posição sênior — a pessoa precisa tomar decisão sob pressão em incidente e evoluir a observabilidade da plataforma.


Responsabilidades
  • Responder a incidentes de produção críticos, conduzindo o diagnóstico até a resolução e o post-mortem.
  • Projetar e manter a stack de observabilidade (métricas, logs, tracing, alertas) das aplicações e serviços.
  • Reduzir MTTR e aumentar a cobertura de monitoramento e a disponibilidade dos serviços.
  • Participar de escala de on-call e melhorar continuamente os runbooks.


Requisitos obrigatórios
  • Experiência sênior em SRE / DevOps / Engenharia de Produção.
  • Domínio de ferramentas de observabilidade (ex.: Datadog, Grafana, Prometheus ou equivalentes).
  • Sólida experiência com ambientes cloud e infraestrutura em produção.
  • Prática consolidada em gestão de incidentes e resposta a on-call.
  • Scripting para automação (Python, Bash ou similar).


Diferenciais
  • Certificações cloud (AWS/GCP/Azure).
  • Experiência com Kubernetes e Infrastructure as Code.
  • Vivência em ambientes de dados / pipelines de crédito.


    CIAL provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.  

Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Software Development

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified