At Tether, we’re pioneering a global financial revolution with cutting-edge blockchain solutions, including the world’s most trusted stablecoin (USDT) and innovative products in energy, data, education, and AI. As a member of our AI model team, you will drive innovation in model serving and inference architectures, optimizing deployment for high performance, scalability, and efficiency across diverse environments—from resource-constrained devices to complex multi-modal systems. This role involves hands-on research, development, and optimization of inference pipelines to deliver low-latency, high-throughput, and memory-efficient AI performance in real-world applications.
Responsibilities
- Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage, ensuring efficiency across diverse environments, including resource-constrained devices and edge platforms.
- Establish clear performance targets such as reduced latency, improved token response, and minimized memory footprint.
- Build, run, and monitor controlled inference tests in both simulated and live production environments.
- Track key performance indicators such as response latency, throughput, memory consumption, and error rates, with special attention to metrics specific to resource-constrained devices.
- Document iterative results and compare outcomes against established benchmarks to validate performance across platforms.
- Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges, specifically those encountered on low-resource devices.
- Set measurable criteria to ensure that these resources effectively evaluate model performance, latency, and memory utilization under various operational conditions.
- Analyze computational efficiency and diagnose bottlenecks in the serving pipeline by monitoring both processing and memory metrics.
- Address issues such as suboptimal batch processing, network delays, and high memory usage to optimize the serving infrastructure for scalability and reliability on resource-constrained systems.
- Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines designed for edge and on-device applications.
- Define clear success metrics such as improved real-world performance, low error rates, robust scalability, and optimal memory usage.
- Ensure continuous monitoring and iterative refinements for sustained improvements.
Requirements
- A degree in Computer Science or a related field.
- Ideally a PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
- Must have knowledge of Metal Shading Language (MSL) and be comfortable writing custom compute shaders from scratch.
- Proven experience in low-level kernel optimizations and inference optimization on mobile devices, with measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications, particularly on resource-constrained devices and edge platforms.
- A deep understanding of modern model serving architectures and inference optimization techniques, including state-of-the-art methods for achieving low-latency, high-throughput performance, and efficient memory management in diverse, resource-constrained deployment scenarios.
- Strong expertise in writing GPU kernels for mobile devices (i.e., smartphones) as well as a deep understanding of model serving frameworks and engines.
- Practical experience in developing and deploying end-to-end inference pipelines, from optimizing models for efficient serving to integrating these solutions on resource-constrained devices.
- Demonstrated ability to apply empirical research to overcome challenges in model serving, such as latency optimization, computational bottlenecks, and memory constraints.
- Proficiency in designing robust evaluation frameworks and iterating on optimization strategies to continuously push the boundaries of inference performance and system efficiency.
- Deep understanding of the math and structure behind Diffusion Models and Vision Transformers.
- Understanding of Pruning, Quantization, Flash Attention, KV Cache, Speculative Decoding (Eagle), etc.
- Experience in designing and optimizing high-performance inference engines using techniques like Tensor Parallelism, Pipeline Parallelism, and Expert Parallelism to handle massive models on GPU clusters.
Benefits
- Global remote work opportunities, collaborating with a team of top-tier talent from around the world.
- Opportunity to work on pioneering projects in fintech, AI, energy, data, and education with a leader in the industry.
Additional details
- Recruitment scams have become increasingly common. Candidates should apply only through official channels (Tether’s careers page).
- Verify the recruiter’s identity via their verified LinkedIn profile or by contacting Tether through its official website.
- Tether does not conduct interviews over WhatsApp, Telegram, or SMS; all communication is done through official company emails and platforms.
- All communication from Tether will come from emails ending in @ or @ (exact domains redacted in source).
- Tether will never request payment or financial details during the hiring process. Any such request is a scam and should be reported immediately.
- Originally posted on Himalayas.