AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

Tether Operations Limited

WorldwideremotePosted 11 days ago
Tether Operations Limited logo

Skill Required

AI-Research-and-EngineeringModel-CompressionQuantization-EngineeringInference-OptimizationMachine-Learning-EngineeringAI EngineerMachine LearningBlockchainObservabilityData StructuresDesign PatternsNLPFulltime

Key highlights

  • Global remote work
  • PhD in NLP, Machine Learning, or related field preferred
  • Expertise in Metal Shading Language (MSL) and GPU kernels for mobile devices required
  • Focus on high-throughput, low-latency, and low-memory AI model serving

Role overview

Join Tether to pioneer a global financial revolution by empowering businesses to integrate reserve-backed tokens across blockchains, enabling secure, instant, and low-cost digital transactions. As a member of the AI model team, you will drive innovation in model serving and inference architectures, optimizing deployment for high performance, scalability, and efficiency—from resource-constrained devices to complex multi-modal systems. Collaborate with cross-functional teams to push the boundaries of AI in real-world applications, ensuring high-throughput, low-latency, and low-memory solutions that deliver tangible value.

Responsibilities

  • Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage.
  • Ensure these pipelines run efficiently across diverse environments, including resource-constrained devices and edge platforms.
  • Establish clear performance targets such as reduced latency, improved token response, and minimized memory footprint.
  • Build, run, and monitor controlled inference tests in both simulated and live production environments.
  • Track key performance indicators such as response latency, throughput, memory consumption, and error rates, with special attention to metrics specific to resource-constrained devices.
  • Document iterative results and compare outcomes against established benchmarks to validate performance across platforms.
  • Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges, specifically those encountered on low-resource devices.
  • Set measurable criteria to ensure that these resources effectively evaluate model performance, latency, and memory utilization under various operational conditions.
  • Analyze computational efficiency and diagnose bottlenecks in the serving pipeline by monitoring both processing and memory metrics.
  • Address issues such as suboptimal batch processing, network delays, and high memory usage to optimize the serving infrastructure for scalability and reliability on resource-constrained systems.
  • Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines designed for edge and on-device applications.
  • Define clear success metrics such as improved real-world performance, low error rates, robust scalability, optimal memory usage and ensure continuous monitoring and iterative refinements for sustained improvements.

Requirements

  • A degree in Computer Science or related field.
  • Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
  • Must have knowledge of Metal Shading Language (MSL). You should be comfortable writing custom compute shaders from scratch.
  • Proven experience in low-level kernel optimizations and inference optimization on mobile devices is essential.
  • Your contributions should have led to measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications, particularly on resource-constrained devices and edge platforms.
  • A deep understanding of modern model serving architectures and inference optimization techniques is required.
  • This includes state-of-the-art methods for achieving low-latency, high-throughput performance, and efficient memory management in diverse, resource-constrained deployment scenarios.
  • Must have strong expertise in writing GPU kernels for mobile devices (i.e., smartphones) as well as a deep understanding of model serving frameworks and engines.
  • Practical experience in developing and deploying end-to-end inference pipelines, from optimizing models for efficient serving to integrating these solutions on resource-constrained devices is required.
  • Demonstrated ability to apply empirical research to overcome challenges in model serving, such as latency optimization, computational bottlenecks, and memory constraints.
  • You should be proficient in designing robust evaluation frameworks and iterating on optimization strategies to continuously push the boundaries of inference performance and system efficiency.
  • Distributed Inference Systems: Designing and optimizing high-performance inference engines using techniques like Tensor Parallelism, Pipeline Parallelism, and Expert Parallelism to handle massive models on GPU clusters.
  • Deep understanding of the math and structure behind Diffusion Models and Vision Transformers
  • Understanding of Pruning, Quantization, Flash attention, KV Cache, Speculative Decoding (Eagle) etc.
  • Excellent English communication skills

Benefits

  • Global remote work opportunities, collaborating with top talent from around the world.
  • Opportunity to work on cutting-edge projects in fintech, AI, energy, data, and education, shaping the future of digital finance and technology.

Additional details

  • Tether’s product suite includes USDT (the world’s most trusted stablecoin), digital asset tokenization services, sustainable energy solutions for Bitcoin mining, AI and peer-to-peer technology breakthroughs (e.g., KEET app), and digital learning initiatives.
  • Tether is a fast-growing, lean industry leader in the fintech space.
  • Recruitment scams have become increasingly common. Candidates should: apply only through official channels (careers page), verify recruiter identities via LinkedIn or the company website, avoid unusual communication methods (e.g., WhatsApp, Telegram, SMS), check for official email domains (@ or @), and never share financial details or payments. Report suspicious activity via the official website.
Apply now