AI Research Engineer (Model Serving & Inference)

Tether Operations Limited, launched in 2014, is a blockchain-enabled platform designed to facilitate the use of fiat currencies in a digital manner.

Tether Operations Limited

Employee count: 51-200

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

As a member of our AI model team, you will drive innovation in model serving and inference architectures for advanced AI systems. Your work will focus on optimizing model deployment and inference strategies to deliver highly responsive, efficient, and scalable performance across real-world applications. You will work on a wide spectrum of systems, ranging from resource-efficient models designed for limited hardware environments to complex, multi-modal architectures that integrate data such as text, images, and audio.

We expect you to have deep expertise in designing and optimizing model serving pipelines and inference frameworks as well as a strong background in advanced model architectures. You will adopt a hands-on, research-driven approach to develop, test, and implement novel serving strategies and inference algorithms. Your responsibilities include engineering robust inference pipelines, establishing comprehensive performance metrics, and identifying and resolving bottlenecks in production environments. The ultimate goal is to enable high-throughput, low-latency, low-memory footprint, and scalable AI performance that delivers tangible value in dynamic, real-world scenarios.

Responsibilities:

Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage. Ensure these pipelines run efficiently across diverse environments, including resource-constrained devices and edge platforms. Establish clear performance targets such as reduced latency, improved token response, and minimized memory footprint.
Build, run, and monitor controlled inference tests in both simulated and live production environments. Track key performance indicators such as response latency, throughput, memory consumption, and error rates, with special attention to metrics specific to resource-constrained devices. Document iterative results and compare outcomes against established benchmarks to validate performance across platforms.
Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges, specifically those encountered on low-resource devices. Set measurable criteria to ensure that these resources effectively evaluate model performance, latency, and memory utilization under various operational conditions.
Analyze computational efficiency and diagnose bottlenecks in the serving pipeline by monitoring both processing and memory metrics. Address issues such as suboptimal batch processing, network delays, and high memory usage to optimize the serving infrastructure for scalability and reliability on resource-constrained systems.
Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines designed for edge and on-device applications. Define clear success metrics such as improved real-world performance, low error rates, robust scalability, optimal memory usage and ensure continuous monitoring and iterative refinements for sustained improvements.

Requirements

A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
Proven experience in large-scale model serving and inference optimization is essential. Your contributions should have led to measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications, particularly on resource-constrained devices and edge platforms.
A deep understanding of modern model serving architectures and inference optimization techniques is required. This includes state-of-the-art methods for achieving low-latency, high-throughput performance, and efficient memory management in diverse, resource-constrained deployment scenarios.
Must have strong expertise in one or more of the following: C/C++, Triton, ThunderKittens, or native CUDA, as well as a deep understanding of model serving frameworks and engines. Practical experience in developing and deploying end-to-end inference pipelines, from optimizing models for efficient serving to integrating these solutions on resource-constrained devices is required.
Demonstrated ability to apply empirical research to overcome challenges in model serving, such as latency optimization, computational bottlenecks, and memory constraints. You should be proficient in designing robust evaluation frameworks and iterating on optimization strategies to continuously push the boundaries of inference performance and system efficiency.

Apply now

Please let Tether Operations Limited know you found this job on Himalayas. This helps us grow!

Apply now

About the job

Apply before

Aug 14, 2025

Posted on

Jun 15, 2025

Job type

Full Time

Experience level

Mid-level

Location requirements

Open to candidates from all countries.

Hiring timezones

Worldwide

Job categories

AI Inference Engineer Mid Level AI Inference Engineer

Skills

AI Research Model Serving Platforms Optimized Inference Application Blockchain Technology Blockchain Mining AI NLP Machine Learning C C++Triton ThunderKittens CUDA Blockchain

About Tether Operations Limited

Learn more about Tether Operations Limited and their company culture.

View company profile

Tether Operations Limited, launched in 2014, is a blockchain-enabled platform designed to facilitate the use of fiat currencies in a digital manner. Since its inception, Tether has aimed at disrupting traditional financial systems by offering a more modern approach to monetary transactions. This enables users to confidently transact with traditional currencies across multiple blockchain networks.

Tether has established itself as the first and largest stablecoin provider globally, offering unmatched liquidity and facilitating billions of transactions. The company operates on a fully transparent basis and ensures that all Tether tokens, primarily the USD₮, are backed 100% by a reserve of corresponding fiat currencies. This commitment to transparency is accompanied by regular disclosures regarding total assets and reserves, reinforcing trust within the cryptocurrency ecosystem.

Apply now

Please let Tether Operations Limited know you found this job on Himalayas. This helps us grow!

Apply now

About the job

Apply before

Aug 14, 2025

Posted on

Jun 15, 2025

Job type

Full Time

Experience level

Mid-level

Location requirements

Open to candidates from all countries.

Hiring timezones

Worldwide

Job categories

AI Inference Engineer Mid Level AI Inference Engineer

Skills

AI Research Model Serving Platforms Optimized Inference Application Blockchain Technology Blockchain Mining AI NLP Machine Learning C C++Triton ThunderKittens CUDA Blockchain

Claim this profile Claim this profile

Tether Operations Limited

Company size

51-200 employees

Founded in

2014

Chief executive officer

Paolo Ardoino

Markets

Blockchain Cryptocurrency Stablecoins Financial Technology (FinTech)Digital Payments Asset Management Decentralized Finance (DeFi)Cross Border Payments Transparency Solutions Reserve Backed Digital Assets

Employees live in

United States

View company profile

Similar remote jobs

Here are other jobs you might want to apply for.

View all remote jobs

Machine Learning Engineer

Gate.io

Employee count: 201-500

Full Time

Machine Learning Engineer

Research Engineer (Cryptography/Security)

MachineFi Lab

Employee count: 51-200

Full Time

Research Engineer

Senior Data Engineer

Biconomy

Employee count: 51-200

Full Time

Senior Data Engineer

Backend Engineer (Remote, Global)

Runloop

Employee count: 1-10

Full Time

Backend Engineer

Research Scientist (Test Time Compute)

Naptha AI

Employee count: 11-50

Full Time

Research Scientist

Infrastructure Engineer

Paradex

Employee count: 11-50

Full Time

Infrastructure Engineer

83 remote jobs at Tether Operations Limited

Explore the variety of open remote roles at Tether Operations Limited, offering flexible work options across multiple disciplines and skill levels.

View all jobs at Tether Operations Limited

Senior Manual QA Engineer (100% remote)

Tether Operations Limited

Employee count: 51-200

Full Time

Senior QA Engineer

Brazil only

AI Research Engineer (Model Evaluation - 100% remote Brazil)

Tether Operations Limited

Employee count: 51-200

Full Time

ML Research Engineer

Spain only

Senior Manual QA Engineer (100% remote - Spain)

Tether Operations Limited

Employee count: 51-200

Full Time

Senior QA Engineer

Nodejs Senior Software Engineer

Tether Operations Limited

Employee count: 51-200

Full Time

Node.Js Backend Engineer

Argentina only

AI Research Engineer (Model Evaluation - 100% remote Argentina)

Tether Operations Limited

Employee count: 51-200

Full Time

ML Research Engineer

United Arab Emirates only

AI Research Engineer (Model Serving & Inference - 100% Remote UAE)

Tether Operations Limited

Employee count: 51-200

Full Time

Mid Level AI Inference Engineer

Top remote companies

Remote companies like Tether Operations Limited

Find your next opportunity by exploring profiles of companies that are similar to Tether Operations Limited. Compare culture, benefits, and job openings on Himalayas.

View all companies

ZH5 jobs

Zero Hash

Tech stack

Zero Hash is a B2B embedded infrastructure platform that allows any platform to integrate digital assets natively into its own customer experience quickly and easily, handling the entire back-end complexity and regulatory licensing.

Fintech Cryptocurrency

PA25 jobs

Paxos

Salaries Benefits Tech stack

Paxos Trust Company is a regulated blockchain infrastructure platform that enables enterprises to tokenize, custody, trade, and settle assets. Their products aim to create a more open and efficient financial system.

Blockchain Digital Assets

ST7 jobs

stakefish

stakefish is a leading validator for Proof of Stake blockchains, offering non-custodial solutions for secure staking.

Blockchain Cryptocurrency

OK14 jobs

OKX

Salaries Tech stack

OKX is a global cryptocurrency exchange and Web3 technology company, offering trading, wallet services, and access to decentralized finance. Founded in 2017, it serves millions of users in over 100 countries.

Cryptocurrency Blockchain Technology

GA76 jobs

Gate.io

Salaries

Gate.io is a leading cryptocurrency exchange founded in 2013, noted for its security and extensive array of trading options.

Cryptocurrency Exchange Blockchain Technology

GE8 jobs

Gemini

Salaries Benefits

Gemini Trust Company, LLC is a cryptocurrency exchange and custodian founded in 2014 by Cameron and Tyler Winklevoss, offering a platform for buying, selling, storing, and earning digital assets. It emphasizes security and regulatory compliance, being a New York trust company.

Digital Payments Digital Asset Trading

Top remote companies

Remote companies like Tether Operations Limited

Find your next opportunity by exploring profiles of companies that are similar to Tether Operations Limited. Compare culture, benefits, and job openings on Himalayas.

View all companies

Find your dream job

Sign up now and join over 85,000 remote workers who receive personalized job alerts, curated job matches, and more for free!

Find your dream job

Sign up now and join over 85,000 remote workers who receive personalized job alerts, curated job matches, and more for free!