AI/LLM Evaluation & Alignment Software Engineer

Save

LeoTech

Salary: 135k-160k USD

United States only

Stay safe on Himalayas

Never send money to companies. Jobs on Himalayas will never require payment from applicants.

At LeoTech, we are passionate about building software that solves real-world problems in the Public Safety sector. Our software has been used to help the fight against continuing criminal enterprises, drug trafficking organizations, identifying financial fraud, disrupting sex and human trafficking rings and focusing on mental health matters to name a few.

Role

This is a remote, WFH role.
As an AI/LLM Evaluation & Alignment Engineer on our Data Science team, you will play a critical role in ensuring that our Large Language Model (LLM) and Agentic AI solutions are accurate, safe, and aligned with the unique requirements of public safety and law enforcement workflows. You will design and implement evaluation frameworks, guardrails, and bias-mitigation strategies that give our customers confidence in the reliability and ethical use of our AI systems. This is an individual contributor (IC) role that combines hands-on technical engineering with a focus on responsible AI deployment. You will work closely with AI engineers, product managers, and DevOps teams to establish standards for evaluation, design test harnesses for generative models, and operationalize quality assurance processes across our AI stack.

Core Responsibilities

Build and maintain evaluation frameworks for LLMs and generative AI systems tailored to public safety and intelligence use cases.
Design guardrails and alignment strategies to minimize bias, toxicity, hallucinations, and other ethical risks in production workflows.
Partner with AI engineers and data scientists to define online and offline evaluation metrics (e.g., model drifts, data drifts, factual accuracy, consistency, safety, interpretability).
Implement continuous evaluation pipelines for AI models, integrated into CI/CD and production monitoring systems.
Collaborate with stakeholders to stress test models against edge cases, adversarial prompts, and sensitive data scenarios.
Research and integrate third-party evaluation frameworks and solutions; adapt them to our regulated, high-stakes environment.
Work with product and customer-facing teams to ensure explainability, transparency, and auditability of AI outputs.
Provide technical leadership in responsible AI practices, influencing standards across the organization.
Contribute to DevOps/MLOps workflows for deployment, monitoring, and scaling of AI evaluation and guardrail systems (experience with Kubernetes is a plus).
Document best practices and findings, and share knowledge across teams to foster a culture of responsible AI innovation.

What We Value

Bachelor's or Master's in Computer Science, Artificial Intelligence, Data Science, or related field.
3–5+ years of hands-on experience in ML/AI engineering, with at least 2 years working directly on LLM evaluation, QA, or safety.
Strong familiarity with evaluation techniques for generative AI: human-in-the-loop evaluation, automated metrics, adversarial testing, red-teaming.
Experience with bias detection, fairness approaches, and responsible AI design.
Knowledge of LLM observability, monitoring, and guardrail frameworks e.g Langfuse, Langsmith
Proficiency with Python and modern AI/ML/LLM/Agentic AI libraries (LangGraph, Strands Agents, Pydantic AI, LangChain, HuggingFace, PyTorch, LlamaIndex).
Experience integrating evaluations into DevOps/MLOps pipelines, preferably with Kubernetes, Terraform, ArgoCD, or GitHub Actions.
Understanding of cloud AI platforms (AWS, Azure) and deployment best practices.
Strong problem-solving skills, with the ability to design practical evaluation systems for real-world, high-stakes scenarios.
Excellent communication skills to translate technical risks and evaluation results into insights for both technical and non-technical stakeholders.

Technologies We Use

Cloud & Infrastructure: AWS (Bedrock, SageMaker, Lambda), Azure AI, Kubernetes (EKS), Terraform, ArgoCD.
LLMs & Evaluation: HuggingFace, OpenAI API, Anthropic, LangChain, LlamaIndex, Ragas, DeepEval, OpenAI Evals.
Observability & Guardrails: Langfuse, GuardrailsAI.
Backend & Data: Python (primary), ElasticSearch, Kafka, Airflow.
DevOps & Automation: GitHub Actions, CodePipeline.

What You Can Expect

Work from home opportunity
Enjoy great team camaraderie.
Thrive on the fast pace and challenging problems to solve.
Modern technologies and tools.
Continuous learning environment.
Opportunity to communicate and work with people of all technical levels in a team environment.
Grow as you are given feedback and incorporate it into your work.
Be part of a self-managing team that enjoys support and direction when required.
3 weeks of paid vacation – out the gate!!
Competitive Salary.
Generous medical, dental, and vision plans.
Sick, and paid holidays are offered.

LeoTech is an equal opportunity employer and does not discriminate on the basis of any legally protected status.

Apply now

Please let LeoTech know you found this job on Himalayas. This helps us grow!

Apply now

About the job

Apply before

Nov 17, 2025

Posted on

Sep 18, 2025

Job type

Full Time

Experience level

Mid-level

Salary

Salary: 135k-160k USD

Location requirements

United States

Hiring timezones

United States +/- 0 hours

About LeoTech

Learn more about LeoTech and their company culture.

View company profile

Apply now

Please let LeoTech know you found this job on Himalayas. This helps us grow!

Apply now

About the job

Apply before

Nov 17, 2025

Posted on

Sep 18, 2025

Job type

Full Time

Experience level

Mid-level

Salary

Salary: 135k-160k USD

Location requirements

United States

Hiring timezones

United States +/- 0 hours

Claim this profile LE

LeoTech

View company profile

Similar remote jobs

Here are other jobs you might want to apply for.

View all remote jobs

United States only

abridge

Employee count: 51-200

Salary: 162k-234k USD

Full Time

GenAI Software Engineer

United States only

Senior Software Engineer - Agentic AI

Code Metal

Employee count: 11-50

Full Time

Find your next opportunity by exploring profiles of companies that are similar to LeoTech. Compare culture, benefits, and job openings on Himalayas.

View all companies

Find your dream job

Sign up now and join over 100,000 remote workers who receive personalized job alerts, curated job matches, and more for free!

Find your dream job

Sign up now and join over 100,000 remote workers who receive personalized job alerts, curated job matches, and more for free!

AI/LLM Evaluation & Alignment Software Engineer

Role

Core Responsibilities

What We Value

Technologies We Use

What You Can Expect

Apply now

About the job

Apply before

Posted on

Job type

Experience level

Salary

Location requirements

Hiring timezones

Job categories

Skills

About LeoTech

Apply now

About the job

Apply before

Posted on

Job type

Experience level

Salary

Location requirements

Hiring timezones

Job categories

Skills

LeoTech

Similar remote jobs

AI / ML Engineer

AI Enablement Engineer

AI/ML Engineer

AI Developer

Software Engineer, Gen AI Platform

Senior Software Engineer - Agentic AI

5 remote jobs at LeoTech

Senior Back End Engineer - Platform

Transcription Software Engineer

AI/NLP Engineer

Site Reliability Engineer

Remote companies like LeoTech

Find your dream job

Find your dream job

Find your dream job

Senior Back End Engineer - Platform

Transcription Software Engineer

AI/NLP Engineer

Site Reliability Engineer

AI / ML Engineer

AI Enablement Engineer

AI/ML Engineer

AI Developer

Software Engineer, Gen AI Platform

Senior Software Engineer - Agentic AI