About Cognitive:
Cognitive is an IT and software engineering services company dedicated to elevating the quality, speed and delivery of today’s US government healthcare programs. With a wealth of clinical expertise and hands-on experience, our team understands the significant challenges our government clients—and their customers—face every day. This real-world experience equips us to develop IT solutions that seamlessly connect all facets of healthcare delivery. At Cognitive, we’re guiding government agencies to the forefront of technology innovation in healthcare delivery.
Position Overview:
*This position is contingent upon contract award*
The SRE/Release Engineer owns the operational reliability, observability, release coordination, and continuous delivery health of the CMS Drug Data Processing System (DDPS) and Payment Reconciliation System (PRS) production environment. This individual bridges engineering and operations, ensuring production releases are executed safely, monitoring systems deliver real-time operational visibility, and automated processes maintain system resilience and availability.
This is a remote position; however, Cognitive hires only in the following designated U.S. states based on contract and business requirements: VA, DC, MD, TN, FL, AZ, CO, OR, and TX.
Key Responsibilities:
- Coordinate and execute all production releases:validate deployment readiness, support cutovers, perform post-implementation validation, and coordinate rollback execution if needed.
- Maintain real‑time monitoring and observability across program systems using standard tools (e.g., AWS CloudWatch, Splunk, Splunk On‑Call, New Relic, Datadog); ensure key operational metrics and system health are visible to stakeholders in real time.
- Drive reliability improvements: identify recurring issues, automate remediation where feasible, reduce manual operational tasks, and improve system resilience and mean time to recovery (MTTR).
- Maintain alerting and on‑call runbooks; coordinate incident response and escalation paths with cybersecurity and DevSecOps leadership.
- Support disaster recovery readiness: plan and execute DR tests, failover/failback exercises, and validate recovery objectives (RTO/RPO) in partnership with cloud architecture.
- Manage release documentation and communications: maintain release schedules and validation reports; coordinate stakeholder communications before, during, and after releases.
- Provide coverage during peak operational windows andparticipatein the on‑call rotation for production releases and monitoring escalations.
Qualifications:
- Bachelor's degree in Computer Science or related field; 4 or more years in site reliability engineering, DevOps, or release management on production federal systems.
- Hands-on experience with AWS CloudWatch (Logs, Events, Alarms), Splunk, and Splunk On-Call/VictorOps for production observability and incident management.
- Experience with New Relic andDataDog for application performance monitoring.
- Experience coordinating production releases, cutovers, and rollback execution on mission-critical systems.
- Experience developing and maintaining operational runbooks, alerting rules, and SRE playbooks.
- Familiarity with Linux environments (RHEL, CentOS, Amazon Linux 2) and Bash scripting for operational automation.
- Ability to pass CMS and internal required background checks for public trust.
Why Join Us?
- Be part of a mission-driven organization making a difference in healthcare IT.
- Collaborate with innovative and passionate professionals that are there to support you at every turn
- Enjoy a supportive work/life balance with the flexibility of a 100% remote company.
- Benefit from opportunities for growth and development in a dynamic environment.