We’re looking for a Senior Data Analyst to support the enterprise Knowledge Graph initiative.
The project connects large-scale company, individual, and relationship data to support compliance, KYC, credit risk, sanctions screening, beneficial ownership analysis, and corporate structure research.
In this role, you’ll work directly with data engineers, product managers, and analysts to validate data, verify query results, investigate discrepancies, and assess data quality across multiple database technologies.
This is a hands-on role focused on delivering clear evidence, including validated results, documented defects, root-cause analysis, and statistical assessments.
Location: Remote (EU) Engagement: Full-time, long-term contract Start Date: ASAP Language: English Time Zone: Ability to overlap with Eastern US hours
What You’ll Do * Validate large datasets loaded into Databricks, PostgreSQL, and graph databases * Confirm data completeness, accuracy, consistency, and structural integrity * Compare query results across different database technologies * Ensure the same queries produce correct and consistent results across platforms * Identify and investigate data discrepancies and quality issues * Determine whether issues come from ETL pipelines, schema mapping, source data, or database-specific behaviour * Develop validation test cases and define expected query results * Maintain data-quality reports, defect logs, and resolution tracking * Support the assessment of compliance and business use cases * Help determine whether each use case requires a graph database or can be handled using a traditional relational database * Apply statistical methods to assess datasets and benchmark results * Analyse distributions, variance, outliers, sampling quality, and measurement reliability * Document findings clearly for engineering and product teams * Work independently with minimal supervision as part of a cross-functional engineering team
Must-Have Requirements * Strong experience in data analysis, data validation, or data quality roles * Experience validating large and complex datasets * Strong SQL skills * Experience working with PostgreSQL or another relational database * Ability to identify discrepancies and perform root-cause analysis * Experience with ETL pipelines and data transformation processes * Understanding of common data-quality issues during data loading and migration * Experience working with large-scale datasets where manual validation is not sufficient * Practical knowledge of statistical analysis, including: * Distribution analysis * Outlier detection * Variance analysis * Sampling validation * Ability to analyse benchmark results and separate real differences from normal performance variation * Experience working across multiple databases or query technologies * Experience with Databricks, Spark, Delta Lake, or similar distributed data platforms * Strong written communication and documentation skills * Ability to work independently in a remote environment * Experience working within a cross-functional engineering team
Nice-to-Have * Hands-on experience with graph databases such as: * Neo4j * TigerGraph * NebulaGraph * ArangoDB * Apache AGE * Understanding of graph data models, nodes, edges, and relationship structures * Experience with Cypher or other graph query languages * Knowledge of Knowledge Graphs or ontology concepts * Experience with MongoDB * Experience with data visualisation or graph analysis tools * Experience in financial services, compliance, KYC, AML, or risk * Understanding of beneficial ownership, sanctions screening, PEP data, or corporate ownership structures * Experience with large company and entity relationship datasets * Familiarity with Bureau van Dijk, Orbis, or similar data sources
Soft Skills * Strong analytical and problem-solving skills * Excellent attention to detail * Clear and structured communication * Ability to explain data issues to technical and non-technical stakeholders * Independent and proactive working style * Strong ownership and accountability * Comfortable working with changing technologies and requirements * Team-oriented approach
Key Deliverables * Documented validation test cases and expected results * Dataset validation reports for each technology * Cross-platform query comparison results * Root-cause analysis for identified discrepancies * Data-quality defect log and resolution tracking * Statistical analysis of benchmark results * Analytical support for graph database use-case assessment