Generative AI Inference Engineer

Stability AI is a company that develops open-source generative AI models for image, language, audio, video, and 3D, aiming to make AI technology accessible to everyone. Their most well-known product is Stable Diffusion, a text-to-image model.

Stability AI

Employee count: 51-200

United States only

Stay safe on Himalayas

Never send money to companies. Jobs on Himalayas will never require payment from applicants.

Generative AI Inference Engineer

About the role:

We are seeking passionate Machine Learning Engineers to join our Inference team, focusing on the creative applications of generative AI models. The ideal candidate will have substantial experience developing and running inference for multi-modal models. A deep understanding of diffusion model architectures and familiarity with workflow tools like ComfyUI are a big plus. You will be expected to leverage and push the boundaries of state-of-the-art inference optimization techniques for multi-modal generative models. This role offers the opportunity to work alongside top researchers and engineers, utilizing cutting-edge high-performance computing resources to make a significant impact in the rapidly evolving field of generative AI.

Responsibilities:

Lead efforts to drive the design, development of customer-facing multi modal ML inference systems.
Work with the Platform and Inference teams on building inference systems for the next generation of models, where you will work on areas such as optimization, model tuning and deployment.
Partner with leading cloud providers to deliver hosted Stability AI inference solutions.
Be a strategic thought partner for leaders across the organization on driving business impact through machine learning
Be part of the team to bring new Stability models and pipelines into existence
Prototype and productionize inference platform improvements and new features

Qualifications:

7+ years working on productionizing machine learning systems, including inference pipeline development
Expert level knowledge on writing and running python services at scale
5+ years working on python scientific stack, pyTorch and at least one high-performance inference framework (e.g. Triton and TensorRT)
Deep understanding of Diffusion Architecture
Experience profiling and optimizing deep neural networks on Nvidia GPUs, using profiling tools such as NVIDIA Nsight
Experience with python-based image manipulation/encoding/decoding frameworks, such as OpenCV
Experience deploying to cloud orchestration systems such as Kubernetes and cloud providers such as AWS, GCP, and Azure
Experience with Docker
Ability to rapidly prototype solutions and iterate on them with tight product deadlines
Strong communication, collaboration, and documentation skills
Experience with the open-source ML ecosystem (HuggingFace, W&B, etc.)

Equal Employment Opportunity:

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.

Apply now

Please let Stability AI know you found this job on Himalayas. This helps us grow!

Apply now

About the job

Apply before

Jul 02, 2026

Posted on

May 03, 2026

Job type

Full Time

Experience level

Senior

Experience

7 years minimum

Location requirements

United States

Hiring timezones

United States +/- 0 hours

Browse similar jobs

Remote Senior Machine-Learning-Engineering Jobs Remote Full Time Machine-Learning-Engineering Jobs Remote Senior Machine-Learning-Engineering Jobs in United States Remote Full Time Jobs in United States Remote Machine-Learning-Engineering Jobs in United States

About Stability AI

Learn more about Stability AI and their company culture.

View company profile

Stability AI is a company focused on designing and implementing solutions in the artificial intelligence (AI) space, with a mission to 'activate humanity's potential' by making foundational AI technology accessible to everyone. Many customers and developers are looking for ways to harness the power of generative AI for creative, analytical, and practical applications. Stability AI addresses this by developing cutting-edge open models across various modalities, including image, video, audio, 3D, and language. This open-access approach is a core differentiator, allowing individuals, researchers, and enterprises to build upon their technology, fostering innovation and creativity without the typical constraints of proprietary software.

The company understands that users, from individual creators to large corporations, face challenges in accessing and utilizing complex AI tools. To solve this, Stability AI provides its models through various channels, including self-hosted memberships, APIs for seamless integration, and cloud platforms for scalability. Their flagship product, Stable Diffusion, a text-to-image model, gained immense popularity for its ability to generate high-quality, detailed images from simple text prompts, empowering users to bring their ideas to life. Beyond image generation, Stability AI is committed to advancing AI in areas like video (Stable Video Diffusion), audio (Stable Audio), and language models (Stable LM), aiming to provide a comprehensive suite of tools for diverse applications. By offering these powerful generative models, Stability AI enables users to streamline workflows, create unique content, and explore new frontiers in their respective fields, whether it's enhancing marketing campaigns, aiding in product design, or facilitating artistic expression.

Tech stack

Learn about the tools and technologies that Stability AI uses to build, market, and sell its products.

View tech stack

Python

JSON

TypeScript

GitHub

TensorFlow

PyTorch

scikit-learn

Grafana

Rippling

Stimulus

Stability AI employees can create an account to update this tech stack.

Apply now

Please let Stability AI know you found this job on Himalayas. This helps us grow!

Apply now