Senior Data Engineer Consultant
DiligenceVault
IndiaremotePosted 1 month ago
Skill Required
Data-EngineeringData-Platform-ArchitectureDatabase-EngineeringData-Architecture-ConsultingAI-ML-Data-InfrastructureData-Engineer-Senior-ConsultantData-Engineering-ConsultantSenior-Data-ConsultantData-Architect-Senior-ConsultantConsultant-Senior-En-Data-EngineeringData Engineerdata engineeringElasticsearchObservabilityBusiness AnalysisSQL ServerEngineeringPostgreSQLSalesforceanalyticalautomationanalyticsPythondesignRabbitMQOpenAPI.NETSparkAzurejQueryAPIsdbtETLSQLandGenerative AIContract
Key highlights
- Contract role (3–6 months + advisory)
- 8-10+ years data engineering/platform architecture required
- Remote with US time zone overlap
- PostgreSQL migration expertise required
- Multi-tenant data residency architecture focus
- Global financial platform impact
Role overview
DiligenceVault is an enterprise B2B SaaS platform digitizing and automating the end-to-end due diligence lifecycle for institutional investors, asset managers, and related stakeholders. The platform processes large volumes of structured and unstructured data from diverse sources and requires a Senior Data Engineering Consultant to help architect a deliberate, AI-native data platform. The consultant will educate leadership, assess current infrastructure, design target architectures (including PostgreSQL migration, canonical data models, data residency, and governance), and produce reference materials for long-term use. The role is a 3–6 month contract with ongoing advisory, fully remote with overlap with US working hours.
Responsibilities
- Teach leadership and senior architects the full spectrum of data engineering, covering traditional foundations and AI-native approaches in depth (e.g., ingestion patterns, transformation, data modeling, data quality, entity resolution, orchestration, semantic layers, AI-native versioning).
- Assess current data infrastructure end-to-end, including mapping data flows, identifying gaps/technical debt, and producing a landscape assessment with current state, target state, gap analysis, and a prioritized roadmap.
- Design the target data platform architecture across ingestion, transformation, storage, serving, and observability layers.
- Architect a PostgreSQL migration plan, including multi-workload support (pgvector, Citus/pg_analytics, OLAP/OLTP separation), connection pooling (PgBouncer/PgCat), read replica topology, partitioning, and handling SQL Server-specific feature dependencies (stored procedures, Query Store, tempdb).
- Define a phased cutover strategy for PostgreSQL migration, including dual-write/shadow-read validation, query translation, and performance benchmarking.
- Design a canonical data layer to unify entities and schemas across heterogeneous sources (CRM, documents, APIs, filings), including entity resolution, schema alignment, conflict resolution, temporal alignment, and master data store with versioning/auditability.
- Specify where AI-native approaches (embedding-based matching, LLM-assisted semantic mapping) add value vs. traditional deterministic rules in canonical data architecture.
- Design a data residency architecture for a multi-tenant, data-sharing platform, covering owner-anchored residency, cross-region data access, regulatory mapping (GDPR, US, APAC), read-path routing, and interactions with search indexing, AI processing, and analytics.
- Design a data governance framework spanning access control (role/attribute-based), data classification (PII detection, sensitivity tagging), lineage/auditability, retention/lifecycle management, consent/data rights, quality accountability, and security controls (encryption, key management, network isolation).
- Define and prioritize use cases for the data platform: cross-source intelligence, customer behavioral insights, automated data enrichment, semantic search, compliance signal detection, and analytics/reporting pipelines.
- Produce reference materials: architecture decision records, data flow diagrams, tool evaluation guides, migration runbooks, and training decks for the team’s independent use.
- Conduct periodic architecture reviews, design consultations, and progress check-ins during execution (ongoing advisory).
Requirements
- 8-10+ years building data platforms across heterogeneous sources at meaningful scale, including foundational architecture (not just individual pipelines).
- Deep expertise in relational databases, specifically PostgreSQL, including hands-on experience with extensions (pgvector, Citus, PostGIS), replication topologies, partitioning, and performance tuning.
- Experience migrating from SQL Server to PostgreSQL is strongly preferred.
- Proven experience designing canonical data models that reconcile entities and schemas across multiple disparate sources, including entity resolution, master data management, and conflict resolution at scale.
- Experience architecting multi-region or data-residency-compliant systems, ideally in a multi-tenant SaaS context with cross-jurisdictional data sharing.
- Strong understanding of data governance: access control models, data classification, lineage, retention policies, and regulatory compliance (GDPR minimum, broader preferred).
- Deep knowledge of traditional data engineering (dimensional modeling, ETL/ELT, CDC, orchestration, query optimization) combined with active engagement with AI-native approaches (ML-driven quality, semantic matching, embedding pipelines, LLM-assisted development).
- Architecture-level thinking: ability to design multi-layer platforms and make defensible technology choices considering scale, cost, team size, and maintainability.
- Exceptional communication and teaching ability: must explain complex concepts clearly and calibrate depth to senior architects and leadership.
- Experience defining data use cases tied to business outcomes, not just building infrastructure in isolation.
Nice to have
- Experience with DiligenceVault’s broader stack or close equivalents: Python/Celery, Elasticsearch, Azure cloud services, .NET APIs, Kestra or similar orchestration.
- Hands-on experience with document processing pipelines: OCR, layout-aware extraction, table parsing from financial PDFs/Word documents.
- Familiarity with modern data stack tools (dbt, Airflow/Dagster, Airbyte/dlt, Snowflake/Databricks) and ability to evaluate them in context.
- Background in financial services, investment management, or due diligence workflows.
- Experience with AI/ML data infrastructure: vector databases, RAG pipelines, feature stores, embedding workflows.
- Track record of producing technical documentation and training materials that teams use independently after the consultant leaves.
Benefits
- Global Influence: Design for a platform used by the world's largest financial institutions in 150+ countries.
- Startup Energy: Flat hierarchy with quick shipping and direct impact.
- Remote Flexibility: Work from anywhere in India while collaborating with a global team.
Additional details
- Job title: Senior Data Engineering Consultant - Platform Architecture & AI-Native Data Strategy.
- Engagement type: Contract / Consulting (3–6 months with ongoing advisory).
- Location: Remote (reasonable 3-4 hours overlap with US working hours required).
- Time commitment: 4-5 hours during initial phase with availability during US working hours (6-11pm IST) required; may be extended if needed.
- Engagement structure: Phase 1 (Education, assessment, roadmap), Phase 2 (Architecture design, use cases, reference materials), Ongoing advisory.
- This is not a staff augmentation role; primary deliverables are knowledge transfer, architectural guidance, and strategic documents, not production code.
- Financial services experience is not required but reduces ramp-up time.
- Originally posted on Himalayas.