Posted today · be early
AI Research Engineer (Multi-Modal & Vision)
Tether Operations Limited
WorldwideremotePosted 1 day ago
Skill Required
AI-Research-EngineerComputer-Vision-EngineerMultimodal-AI-EngineerMachine-Learning-EngineerVision-Language-Model-EngineerAI-ML-Research-EngineerML-AI-Research-EngineerSenior-AI-Research-EngineerML-Research-EngineerMachine-Learning-Research-EngineerSenior-AI-Vision-ResearcherAI-ML-Research-ScientistResearch-EngineerAI EngineerAIMachine Learningrelated fieldEngineeringBlockchainbuildingfintechetc.)designGitRustandFulltime
Key highlights
- Fully remote position
- Focus on vision-language models
- Requires MS/PhD degree
- Requires top-tier AI conference publications
- Global fintech company
Role overview
As a member of the AI model team, you will drive innovation in training and optimizing vision-language models with a focus on real-world deployment. Your work will span the full model development lifecycle—from data curation and training pipeline design to model evaluation and optimization—to build models that are both highly capable and practical for production environments.
Responsibilities
- Conduct end-to-end research and engineering on vision-language models, covering training, evaluation, and optimization across the full model development lifecycle.
- Design and implement post-training pipelines including supervised fine-tuning, knowledge distillation, and reinforcement learning from human feedback.
- Develop and maintain high-quality multimodal datasets, including data curation, filtering, and balancing for domain-specific tasks.
- Drive model efficiency and deployability, adapting models for resource-constrained environments using compression and optimization techniques.
- Design and implement evaluation frameworks and benchmarks to measure model performance, robustness, and real-world task success.
- Build and scale training workflows across distributed GPU infrastructure.
- Identify and resolve bottlenecks in training pipelines to achieve state-of-the-art model quality on target benchmarks.
- Contribute to and leverage open-source ecosystems including models, datasets, and tooling to accelerate development.
- Stay current with the latest research in multimodal learning and vision-language systems, translating relevant findings into practical improvements.
- Publish research findings in top-tier AI conferences and journals where applicable.
Requirements
- Degree in Computer Science, Machine Learning, or a related field (MS/PhD preferred).
- Strong experience with multimodal post-training workflows including supervised fine-tuning, knowledge distillation, and reinforcement learning from feedback.
- Hands-on experience with parameter-efficient fine-tuning and distributed training frameworks.
- Demonstrated ability to build and improve vision-language models with measurable results on standard benchmarks or real-world tasks.
- Experience adapting models for resource-constrained environments.
- Proven open-source contributions in multimodal AI on GitHub or HuggingFace.
- Publications at top AI conferences (NeurIPS, ICML, ICLR, CVPR, ECCV etc.)
- Excellent English communication skills.
Benefits
- Work remotely from every corner of the world.
- Collaborate with some of the brightest minds in the fintech space.
Additional details
- Tether operates a product suite including USDT, Tether Power (energy/Bitcoin mining solutions), Tether Data (AI and P2P tech), and Tether Education.
- Recruitment Scam Warning: Apply only through official channels; Tether does not use third-party agencies, does not interview via WhatsApp/Telegram/SMS, and will never request payment.
- Recruitment communication comes only from official company emails.
- Originally posted on Himalayas.