Staff AI Infra Engineer (serving API)

Company: Genmo Replay
Location: San Francisco
Posted on: November 6, 2024

Job Description:

We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Join us in shaping the future of AI and pushing the boundaries of what's possible in video generation.Role OverviewWe are looking for a senior/staff software engineer to join our inference team. In this role, you will be responsible for designing and scaling our inference systems as they grow to support over millions of users across more than 20 different data centers.Key Responsibilities

Develop high-performance, high-throughput, efficient, and low-latency inference pipelines.
Design, develop, and maintain scalable backend services that support our AI-powered content creation platform.
Implement and optimize model serving infrastructure using Kubernetes and other cloud-native technologies.
Collaborate with ML engineers to transition models from research to production.
Design APIs for integrating our AI capabilities into our partner ecosystem.
Implement monitoring, logging, and alerting systems for backend services and model inference.
Develop monitoring infrastructure for our ML serving pipeline and apply advanced model compression and optimization techniques (quantization, pruning, distillation) to improve inference performance.Qualifications
Bachelor's or Master's degree in Computer Science, Software Engineering, or a related field.
5+ years of experience in software engineering, with at least 3 years focusing on backend systems and ML infrastructure.
Must Have:

Strong past experience with Ray or Kubernetes.
Strong proficiency in Python and at least one systems programming language (Rust, C++ or Go).
Solid understanding of model serving frameworks (e.g., TensorFlow Serving, NVIDIA Triton).
Experience with a ML framework such as TensorFlow, PyTorch, or JAX.
Experience with model compression and optimization techniques.
Strong knowledge of cloud platforms (AWS, GCP, or Azure) and their ML-specific services.
Familiarity with distributed systems and microservices architectures.
Experience with high-performance, low-latency systems.
Ideal candidate will have:
- Experience with GPU programming is a plus.Additional InformationThe role is based in the Bay Area (San Francisco). Candidates are expected to be located near the Bay Area or open to relocation.Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company.
  #J-18808-Ljbffr

Keywords: Genmo Replay, Rohnert Park , Staff AI Infra Engineer (serving API), Engineering , San Francisco, California

Click here to apply!

Didn't find what you're looking for? Search again!

Let San Francisco recruiters find you. Post your resume for free!

Get San Francisco Engineering jobs via email.

View more Rohnert Park Engineering jobs

Other Engineering Jobs

Customer Engineer II, Applied AI, Google Cloud
Description: MidExperience driving progress, solving problems, and mentoring more junior team members deeper expertise and applied knowledge within relevant area.Minimum Qualifications: ul li Bachelor's degree (more...)
Company: Google Inc.
Location: San Francisco
Posted on: 11/10/2024

SENIOR BACKEND ENGINEER
Description: Car IQ is driving payments. br br At Car IQ we are on a mission to transform the payment industry, and we're starting with fleet vehicles. As a leading payment provider it is our job to deliver what (more...)
Company: Car IQ
Location: San Francisco
Posted on: 11/10/2024

Frontend Engineer
Description: Join Amazon as a Frontend Engineer to innovate store architecture, build interactive experiences, and enhance global impact on customer shopping. Opportunities in San Francisco, CA.DescriptionAre you (more...)
Company: Teachmecode
Location: San Francisco
Posted on: 11/10/2024

Salary in Rohnert Park, California Area | More details for Rohnert Park, California Jobs |Salary

SR MACHINE LEARNING ENGINEER, INCLUSIVE / RESPONSIBLE AI
Description: About Pinterest br br You can get further details about the nature of this opening, and what is expected from applicants, by reading the below. br br Millions of people across the world come to (more...)
Company: Pinterest
Location: San Francisco
Posted on: 11/10/2024

Backend Engineer San Francisco Mar 3 3:37 AM
Description: Our mission is to accelerate business adoption of electric vehicle fleets. Electric vehicles are superior for just about every business use-case and businesses are adopting EVs quickly. Standard Fleet (more...)
Company: Standardfleet
Location: San Francisco
Posted on: 11/10/2024

Full Stack Engineer
Description: About UsAuditoria is an AI-driven SaaS automation provider for corporate finance that automates back-office business processes involving tasks, analytics, and responses in Vendor Management, Accounts (more...)
Company: Auditoria.AI
Location: San Francisco
Posted on: 11/10/2024

Field Solutions Engineer
Description: Available Locations: Remote MexicoWhat you'll do as a Solutions EngineerYou are the technical lynchpin through the entire sales cycle - pre and post sales. You will work closely with our Enterprise prospects (more...)
Company: CloudFlare
Location: San Francisco
Posted on: 11/10/2024

Senior Backend Engineer, Benefits Marketplace Engineering San Francisco, CA
Description: Senior Backend Engineer, Benefits MarketplaceAbout RipplingRippling gives businesses one place to run HR, IT, and Finance. It brings together all of the workforce systems that are normally scattered across (more...)
Company: Rippling
Location: San Francisco
Posted on: 11/10/2024

Senior Machine Learning Engineer (Modeling), Underwriting and Credit
Description: Senior Machine Learning Engineer Modeling , Underwriting and CreditBlock Published 10 Aug 2024 Share this job San Francisco, CA, USA USD 198000 - USD 297000 annually Full Time Share this jobRole Highlights (more...)
Company: KinderKids Learning Center
Location: San Francisco
Posted on: 11/10/2024

Backend Engineer, Accounting Platform
Description: About Wholesail br is an early stage startup that builds software to connect wholesalers and their buyers by automating and digitizing invoicing and payment in the 5 trillion wholesale trade industry. (more...)
Company: Wholesail
Location: San Francisco
Posted on: 11/10/2024

Loading more jobs...

Staff AI Infra Engineer (serving API)

Didn't find what you're looking for? Search again!

Other Engineering Jobs

Log In or Create An Account