For Employers

Meta

Principal Software Engineer, Systems

Posted a day ago
$271K - $347K per year
10+ years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Define the technical vision and architecture for agentic reliability systems while personally designing and shipping production-grade AI agents. Lead the development of autonomous systems for incident investigation, mitigation, and infrastructure operations at Meta scale.

Meta is seeking an experienced Software Engineer to build the next generation of AI systems for production reliability. This role sits at the intersection of large-scale distributed systems, incident response, and applied AI. You will lead the development of an AI agent that can autonomously investigate production incidents, identify likely root causes, create safe mitigation plans, and execute them while humans supervise and can intervene. You will improve the agent's accuracy, autonomy, and real-world impact, while identifying other opportunities to apply agentic systems across incident prevention, detection, mitigation, observability, and infrastructure operations. This may include improving the agent's architecture, context, and tooling, or adapting foundation models through fine-tuning, reinforcement learning, distillation, pruning, and inference optimization. This is an applied engineering role, not a research position. The ideal candidate combines deep infrastructure expertise with sound judgment about when and how to apply AI to production problems, a long-term technical vision, and the ability to move quickly from an ambitious idea to a production system operating safely at Meta scale. This is also a deeply hands-on senior IC role. You will personally prototype, build, evaluate, and ship AI systems, working directly in the code and with production data. A central challenge is creating systems that improve from experience: learning from outcomes, identifying their own failure modes, testing changes, and safely increasing their effectiveness over time.

Responsibilities

  • Define the technical vision and architecture for agentic reliability systems across Meta
  • Personally design, code, and ship production agentic systems for complex infrastructure problems
  • Lead the development of an AI agent for production incident investigation and mitigation, advancing it toward accurate, trusted, and safely supervised autonomous action
  • Develop major improvements in agent reasoning, context, tool use, planning, evaluation, learning, and safe execution
  • Explore and apply techniques including automated hill climbing, fine-tuning, reinforcement learning, model routing, distillation, pruning, and inference optimization, selecting the simplest approach that produces measurable gains
  • Identify new high-value applications of AI across incident prevention, detection, mitigation, observability, and infrastructure operations
  • Move rapidly from ambiguous problems to prototypes, validate them against real production workloads, and develop successful approaches into reliable systems at scale
  • Build evaluation and experimentation systems that connect agent quality to outcomes such as investigation accuracy, successful mitigation, incident duration, and reduced operational work
  • Build closed-loop improvement systems that turn production outcomes into evaluations, experiments, and better agent behavior
  • Establish architectures and guardrails for production actions, including authorization, independent validation, auditability, rollback, and human oversight
  • Partner with Infrastructure, AI, Product, and Reliability leaders to integrate agentic capabilities into Meta’s production ecosystem
  • Influence technical strategy across organizations and mentor other engineers working on distributed systems and applied AI


Minimum Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 12+ years of software engineering experience, including experience building and operating large-scale distributed or infrastructure systems
  • Experience setting technical direction and leading complex, multi-year engineering efforts across organizational boundaries
  • Experience applying AI or machine learning systems to production problems
  • Demonstrated experience moving from technical concept to production deployment and measurable impact
  • Experience diagnosing complex production systems using telemetry, code, configuration, and dependency information
  • Experience coding in languages such as C++, Java, Python, Rust, or equivalent
  • Recent hands-on experience building and shipping complex production systems, with the ability to move directly between architecture, experimentation, debugging, and implementation
  • Experience influencing senior engineers and leaders without direct organizational authority


Preferred Qualifications

  • Experience with autonomous or semi-autonomous production actions and their safety, authorization, and rollback mechanisms
  • Experience building systems that operate at hyperscale under strict reliability and latency requirements
  • Experience improving agent quality through evaluation, context engineering, fine-tuning, reinforcement learning, distillation, or model-serving optimization
  • Track record of identifying unconventional opportunities, rapidly prototyping solutions, and changing the technical direction of a large organization
  • Deep expertise in observability, incident response, change safety, distributed systems, or production infrastructure
  • Experience building self-improving or self-evolving systems, including automated experimentation, feedback loops, hill climbing, reinforcement learning, or recursive self-improvement
  • Experience building production AI agents that reason across telemetry, code, configuration, deployments, and operational knowledge


$271,000/year to $347,000/year + bonus + equity + benefits

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Implementation Specialist (Software)

Full Time Canada Software Development

Cloud Operational Engineer Level 3

Full Time Philippines Software Development

Clinical Developer

Full Time India Software Development

Solution Specialist

Full Time Australia Software Development

Forensic Mechanical Engineer

Full Time United States $130K - $170K per year Software Development

Technical Architect - Azure Certified

Full Time India Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 219,372+ Jobs in Principal Software Engineer

Answer easy questions

Answer easy questions

219,372+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified