For Employers

MARGO

Network Reliability Engineer

Posted 2 hours ago
200 - 250 per hour
2-5 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Build and maintain large-scale AI infrastructure while ensuring system stability, scalability, and security. Collaborate with engineering teams to troubleshoot production incidents and implement observability solutions.

 

#HPC #AI #GPU #CLUSTERS

 

YOUR DAILY ROUTINE

- Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents- Troubleshoot high-impact production issues in collaboration with other engineering teams

- Participate in an on-call rotation to handle incidents and ensure service continuity

- Implement and maintain observability solutions to monitor AI infrastructure and application health

- Contribute to AI infrastructure lifecycle management across different environments and countries

- Promote and apply best practices in terms of stability, resiliency, scalability, and security

- Maintain clear technical documentation for tools and procedures

- Contribute to system and tool evolution based on production feedback

- Collaborate closely with development teams to ensure infrastructure readiness- Participate in team rituals and knowledge-sharing initiatives

 

ABOUT YOU

 

🎯 SOFTSKILLS : 

- Proactive and solution-oriented mindset

- Passion for automation and continuous improvement

- Strong collaboration and communication skills

- Ability to work independently and in a team

- Willingness to mentor and share knowledge

 

💻 HARDSKILLS : 

- Experience with Go or Python 

- Strong scripting skills (Bash, Python)

- Hands-on experience with Linux systems (Ubuntu/Debian)

- Preferred hands-on experience with GPU & HPC infrastructure 

- Knowledge of networking (VLAN/LAN, TCP/IP, DNS, BGP, load-balancing, IPv6, etc.)

- Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.)

- Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.)

- Experience managing relational databases (MariaDB)

- Understanding of CI/CD pipelines (GitLab)

- Comfortable with English (written and spoken)

 

\n


\n
200 zł - 250 zł an hour
\n

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Principal Architect - Enterprise AI

Full Time United States $126K - $244K per year Software Development

L3 Technical Support Engineer

Full Time Worldwide $1100 - $1300 per month Software Development

Frontend Engineer (React.js)

Full Time Spain Software Development

Shopify/Amazon Supply Chain Operations Specialist

Full Time Philippines 70000 - 85000 per month Software Development

Senior Technical Business Analyst

Full Time Oman, Poland, Romania +1 more Software Development

Senior Backend Engineer - Node.js

Full Time United States Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 211,823+ Jobs in Software Development

Answer easy questions

Answer easy questions

211,823+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified