The Operations Associate will manage administrative tasks, travel logistics, and event coordination to support the CEO and Chief of Staff. They will also be responsible for team onboarding, vendor management, and maintaining internal documentation systems.
You will manage administrative operations, including coordinating executive meetings, travel logistics, and team events. Additionally, you will handle team onboarding, expense processing, and the implementation of AI-driven systems to streamline business workflows.
Conduct novel research in Protocol Learning with the goal of publishing in tier-1 ML conferences. Own and solve a specific technical problem that blocks Protocol Learning at scale.
You will architect, build, and scale the infrastructure orchestration and distributed compute platform for decentralized model training. This includes managing multi-cloud deployments, ensuring fault-tolerant distributed ML systems, and handling real-world network conditions.
You will design, build, and ship well-scoped projects on the Agora roadmap while contributing to the production distributed training codebase. You will work hands-on with large-scale training infrastructure and collaborate with engineering and research teams.
You will implement and optimize model-parallel training systems for large models on heterogeneous hardware across low-bandwidth, high-latency networks. Additionally, you will build robust infrastructure for fault tolerance, state synchronization, and performance monitoring across hundreds of devices.
You will own the end-to-end inference stack, including pipeline-parallel execution, routing, and transport mechanisms. Additionally, you will invent and implement novel algorithms to ensure fast, reliable inference on consumer hardware over the public internet.
Identify and solve open research problems related to protocol learning, including communication efficiency and robustness in decentralized environments. Collaborate with the engineering team to implement research findings into live training runs and publish results in top-tier venues.
You will build and manage the end-to-end RL post-training stack, including rollout ingestion, reward computation, and policy updates. Additionally, you will invent new algorithms to handle asynchronous, high-latency, and decentralized training environments.
You will own the threat model for a decentralized training and inference network, identifying risks like model poisoning and data extraction. Additionally, you will design and implement statistical verification algorithms to ensure the integrity of work performed by untrusted nodes.