|
Описание: |
Senior Data Engineer — Data Platform & Orchestration * Work Format: Full-time; remote * Department: Data Science * Reports to: Head of Data Science * Language: English; working Russian is an advantage
About Deep Knowledge Group Deep Knowledge Group (www.dkv.global) builds large-scale analytical databases, investment-intelligence systems, industry ecosystem platforms, competitiveness indices and interactive dashboards across frontier industries including AI, Longevity, HealthTech, DeepTech and QuantumTech.
The Data Science department builds and maintains the industry intelligence databases behind those products — organisations, people, products and the relationships between them, across multiple sectors and regions. Every one of them starts as messy public data and has to end up as a clean, deduplicated, verifiable dataset that analysts and downstream products can rely on. About the Role We currently build each database the way the last one was built. Collectors are source-specific scripts, scheduling is manual, quality is whatever the person who built it checked, and nothing links one database to another. That works at our current size and will not work at the next one.
We are looking for a Senior Data Engineer to own the shared platform that replaces this: one collection and enrichment framework that every database runs on, with orchestration, monitoring, provenance and entity resolution as properties of the system rather than as things individual engineers remember to do.
This is an ownership role. You will work directly with the Head of Data Science on architecture, and you will be the person other engineers bring design decisions to. We are not looking for an extra pair of hands; we are looking for a second point of technical decision-making in the department. Key Responsibilities * Design and own the shared collection and enrichment framework: what is reusable, what stays source-specific, and where the line between them sits. * Own orchestration end to end — scheduling, dependencies, retries, backfills, incremental and idempotent re-runs. * Make entity resolution a permanent service rather than a per-project effort: blocking, matching, thresholds, and a defensible account of why a merge was made. * Design the provenance model: every value traceable to its source, its retrieval date and the logic that transformed it, six months after the fact. * Establish the staging-and-promotion rule and enforce it: nothing writes to production tables directly. * Build monitoring and observability for data, not only for jobs — freshness, coverage, null rates, distribution drift, source schema changes. * Define the quality standard: accuracy, completeness, duplicate rate, freshness, provenance — and the instrument that measures it. * Set the schemas, conventions and review standards that the rest of the department works to. * Design the deterministic backbone of our agent network, and decide where an agent is warranted and where plain code is cheaper, faster and testable. * Migrate existing databases onto the shared framework alongside the engineers who built them. * Mentor and unblock other engineers; take architectural decisions off the Head of Department’s desk.
Required Skills * Strong Python — you write services and pipelines, not notebooks. * You have owned a data platform, not just tasks on one — orchestration, scheduling, monitoring, backfills. * Airflow, Dagster or Prefect in production, with a considered view on why that one. * SQL and PostgreSQL in depth — indexing, query plans, schema design, migrations. * Entity resolution or record linkage across sources that disagree with each other, including how you measured whether it worked. * You have set the standards other people then worked to: schemas, conventions, review. * Docker, CI/CD, and comfort running production systems on plain infrastructure rather than a managed platform. * Practical experience with LLM APIs and with validating non-deterministic output before it reaches a production table. * Data quality discipline: you can state a dataset’s duplicate rate or field staleness with evidence, not by inspection. * The ability to write a design down clearly enough that someone can argue with it.
Strong Advantages * Production collection from public web sources: anti-bot, JS-rendered pages, undocumented APIs, source drift. * Knowledge graphs, ontologies and canonical entity models across multiple domains. * Data lineage and governance tooling in production, used by people other than its author. * Agent orchestration frameworks, and a clear view of their limits. * Vector search and RAG pipelines. * Experience turning a team’s ad-hoc scripts into a framework that team then adopted. * OSINT methodology and source verification.
Відгукнутись на вакансію |