AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide
Tether Operations Limited
WorldwideremotePosted 16 days ago
Skill Required
AI-Research-and-EngineeringModel-CompressionQuantization-EngineeringInference-OptimizationMachine-Learning-EngineeringRemote-Senior-AI-EngineerAI-ML-Research-EngineerAI-Research-EngineerML-AI-Research-EngineerResearch-EngineerAI EngineerAIMachine Learningrelated fieldEngineeringBlockchainObservabilityData StructuresdesigningbuildingfintechdesignDesign PatternsRustNLPandFulltime
Key highlights
- Remote work from anywhere in the world
- PhD in NLP, Machine Learning, or related field preferred
- Expertise in Metal Shading Language (MSL) and GPU kernels for mobile devices required
- Focus on high-performance, low-latency AI model serving and inference optimization
Role overview
Join Tether to pioneer a global financial revolution by empowering businesses to integrate reserve-backed tokens across blockchains, enabling secure, instant, and low-cost digital transactions. As a member of the AI model team, you will drive innovation in model serving and inference architectures, optimizing deployment and inference strategies for high-performance, scalable AI systems across diverse environments, from resource-constrained devices to complex, multi-modal architectures. Your work will focus on delivering high-throughput, low-latency, and memory-efficient AI performance in real-world applications.
Responsibilities
- Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage.
- Ensure these pipelines run efficiently across diverse environments, including resource-constrained devices and edge platforms.
- Establish clear performance targets such as reduced latency, improved token response, and minimized memory footprint.
- Build, run, and monitor controlled inference tests in both simulated and live production environments.
- Track key performance indicators such as response latency, throughput, memory consumption, and error rates, with special attention to metrics specific to resource-constrained devices.
- Document iterative results and compare outcomes against established benchmarks to validate performance across platforms.
- Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges, specifically those encountered on low-resource devices.
- Set measurable criteria to ensure that these resources effectively evaluate model performance, latency, and memory utilization under various operational conditions.
- Analyze computational efficiency and diagnose bottlenecks in the serving pipeline by monitoring both processing and memory metrics.
- Address issues such as suboptimal batch processing, network delays, and high memory usage to optimize the serving infrastructure for scalability and reliability on resource-constrained systems.
- Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines designed for edge and on-device applications.
- Define clear success metrics such as improved real-world performance, low error rates, robust scalability, optimal memory usage and ensure continuous monitoring and iterative refinements for sustained improvements.
Requirements
- A degree in Computer Science or related field.
- Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
- Must have knowledge of Metal Shading Language (MSL). You should be comfortable writing custom compute shaders from scratch.
- Proven experience in low-level kernel optimizations and inference optimization on mobile devices is essential.
- Your contributions should have led to measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications, particularly on resource-constrained devices and edge platforms.
- A deep understanding of modern model serving architectures and inference optimization techniques is required.
- This includes state-of-the-art methods for achieving low-latency, high-throughput performance, and efficient memory management in diverse, resource-constrained deployment scenarios.
- Must have strong expertise in writing GPU kernels for mobile devices (i.e., smartphones) as well as a deep understanding of model serving frameworks and engines.
- Practical experience in developing and deploying end-to-end inference pipelines, from optimizing models for efficient serving to integrating these solutions on resource-constrained devices is required.
- Demonstrated ability to apply empirical research to overcome challenges in model serving, such as latency optimization, computational bottlenecks, and memory constraints.
- You should be proficient in designing robust evaluation frameworks and iterating on optimization strategies to continuously push the boundaries of inference performance and system efficiency.
- Distributed Inference Systems: Designing and optimizing high-performance inference engines using techniques like Tensor Parallelism, Pipeline Parallelism, and Expert Parallelism to handle massive models on GPU clusters.
- Deep understanding of the math and structure behind Diffusion Models and Vision Transformers
- Understanding of Pruning, Quantization, Flash attention, KV Cache, Speculative Decoding (Eagle) etc.
- Excellent English communication skills.
Benefits
- Global talent powerhouse working remotely from every corner of the world.
- Opportunity to collaborate with some of the brightest minds in the fintech space, pushing boundaries and setting new standards.
Additional details
- Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
- Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
- Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
- Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
- Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
- We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
- Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
- Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page:
- Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
- Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
- Double-check email addresses. All communication from us will come from emails ending in @ or @
- We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
- When in doubt, feel free to reach out through our official website.
- Originally posted on Himalayas