Blockchain Data Analyst & Researcher

Lukka

WorldwideremotePosted 1 month ago
Lukka logo

Skill Required

Blockchain-Data-AnalystData-EngineerBlockchain-ResearcherCrypto-AnalystData-ScienceBlockchain-Research-AnalystBlockchain-Analytics-SpecialistCryptocurrency-Data-AnalystBlockchain-AnalystCrypto-Research-AnalystData AnalystBlockchainSystem DesignQuery OptimizationMachine Learningdata engineeringData StructuresSolidityScikit-learnEngineeringWebSocketdebugginganalyticsbuildingSparkPythonPandasdesignFlinkRedisjQueryDeFiAPIsSassNode.jsdbtContract

Key highlights

  • 3+ years hands‑on experience in Data Science or Data Engineering
  • Strong proficiency in Apache Flink or Spark Streaming and Apache Iceberg or Apache Hudi
  • Experience with AWS Data Platform (S3, RDS, Spark SQL, PySpark)
  • Opportunity to work directly on digital‑asset data and DeFi protocols
  • Startup‑style culture with visible impact
  • Inclusive environment that promotes from within

Role overview

As a Blockchain Data Analyst & Researcher, you will be a pivotal force in our data intelligence efforts, serving primarily as a blockchain data protocol researcher. Your core mission will involve a deep dive into the complex internals and mechanics of various blockchain networks, reverse‑engineering sophisticated DeFi protocols, and translating these complex systems into actionable, structured data.

Responsibilities

  • Maintain an AI‑native and blockchain‑forward mindset, proactively leveraging emerging trends and generative technologies to stay at the vanguard of the crypto revolution while maximizing work efficiency and uncovering novel, high‑alpha data opportunities.
  • Develop new derivative data models and datasets for client insights and competitive advantages.
  • Map, cleanse, and normalize raw, disparate blockchain data (transaction logs, smart contract events, state changes) into consistent, structured relational and non‑relational schemes for consumption by analytics, machine learning, and AI tools.
  • Architect intuitive data models turning on‑chain data into elegant, unified abstractions and provide granular, high‑integrity insights across various protocols.
  • Deconstruct DeFi mechanics: reverse‑engineer smart contracts and on‑chain events, translating complex interactions into clean, production‑grade SQL/dbt models.
  • Scale ecosystem coverage: lead the charge in indexing new blockchains, design schemas to transform raw chain data into the internal "source of truth," ensuring every block is captured with 100% accuracy.
  • Engineer signal from noise: define metrics that matter (TVL, Volume, Fees), act as gatekeeper of data quality, identify and filter wash trading, bot activity, and Sybil attacks.
  • Architect high‑throughput ingestion pipelines leveraging external data sources (REST/WebSocket APIs, JSON‑RPC node queries).
  • Apply strong ETL fundamentals to manage schema mismatches between heterogeneous sources and destinations.
  • Systematically debug: reproduce issues, isolate root causes, validate fixes.
  • Work comfortably with large, semi‑structured, or undocumented data sources.

Requirements

  • 3+ years hands‑on experience in Data Science, Data Engineering, or a hybrid role.
  • Strong proficiency with Apache Flink or Spark Streaming for building and operating real‑time data pipelines.
  • Experience with Apache Iceberg or Apache Hudi for managing large‑scale Data Lakehouses with schema evolution, time travel, and ACID guarantees over petabytes of raw blockchain data.
  • Familiarity with AWS Data Platform: S3 for scalable data lake storage; RDS for structured relational workloads; Spark SQL and PySpark for interactive querying and ad‑hoc analysis.
  • Proficiency in high‑throughput caching and distribution technologies such as Hazelcast, Apache Ignite, or Redis for low‑latency data access.
  • Advanced SQL query writing, including window functions, query plan analysis, and performance tuning at scale.
  • Python skills: Pandas, Polars for data wrangling; Scikit‑learn for model development.
  • Ability to architect high‑throughput ingestion pipelines that leverage external data sources (REST/WebSocket APIs, JSON‑RPC node queries).
  • Strong grasp of ETL fundamentals (extract, transform, load) and managing schema mismatches.
  • Systematic debugging mindset: reproducing issues, isolating root causes, validating fixes.
  • Comfortable working with large, semi‑structured, or undocumented data sources.

Nice to have

  • Blockchain or crypto analytics background strongly preferred; familiarity with on‑chain data structures, transaction semantics, and asset classification.

Benefits

  • Opportunity to work directly on data for digital assets, an industry‑leading product with well‑known clients.
  • Working with multiple AI solutions daily.
  • Exposure to both technical problem‑solving and customer‑facing challenges.
  • A growing Poland‑based team with direct collaboration across US and global stakeholders.
  • A startup‑style culture where your impact is visible every day.
  • Ability to own projects and take ideas from inception to delivery.
  • Less red tape and a fast‑moving environment at the forefront of a rapidly growing industry.
  • We invest in teammates, promote from within, and foster an inclusive space for curious minds to thrive and innovate.

Additional details

  • Originally posted on Himalayas
  • We move fast and continue to be ahead of the curve.
  • At Lukka, we strive to anticipate and respond to the needs of this new and ever‑changing ecosystem of digital assets.
  • You could be part of a growing company bridging the gap between business and blockchain.
  • So you’ve scrolled, your interest is piqued and hopefully read all of the important things...now what?
  • If this sounds like you, apply!
Apply now