Data Engineer
B2B Contract 16 500 - 28 000 PLN + VAT
Get to know us better
CodiLime is a software and network engineering industry expert and the first-choice service partner for top global networking hardware providers, software providers and telecoms. We create proofs-of-concept, help our clients build new products, nurture existing ones and provide services in production environments. Our clients include both tech startups and big players in various industries and geographic locations (US, Japan, Israel, Europe).
While no longer a startup - we have 250+ people on board and have been operating since 2011 we’ve kept our people-oriented culture. Our values are simple:
Act to deliver.
Disrupt to grow.
Team up to win.
The project and the team
You will join the team behind a large-scale, centralized data platform built for a global consulting organization. The platform is the shared source of company data behind several of the firm's internal products, and is used daily by consultants for company research, including support for Mergers & Acquisitions (M&A) engagements.
This is fundamentally a data engineering role combined with software engineering: you'll be designing, coding, testing, and operating production Python systems - pipelines, libraries, and services - that move, transform, and serve data at scale. This is not a role focused on configuring tools or writing one-off queries. You'll be building reliable, maintainable software that powers our data platform. The goal is a unified, enterprise-grade dataset of 300M+ company records, integrated from 10+ external and internal sources.
The platform delivers firm-level and site-level data - firmographics, technographics, and hierarchical relationships (parent company, subsidiary, site) - alongside key business metrics such as revenue, CAGR, EBITDA, headcount, M&A activity, competitors, industry classification, and web traffic. Data needs to stay accurate, well-structured, and fast to query as both the dataset and the number of consumers keep growing.
Technology stack:
Languages: Python, SQL
Data platform: Snowflake, dbt
Workflow orchestration: Apache Airflow (complex DAGs), running on Kubernetes
Data processing: Apache Spark on Azure Databricks
Data tooling and DBs: pandas, Polars, PyArrow, DuckDB, PySpark, PostgreSQL, Redis
Cloud: Azure (AKS, Blob Storage, ACR, Databricks, OpenSearch, Azure AI Search)
API & services: FastAPI (REST, async), API Gateway
Testing & code quality: pytest, mypy/pyright, ruff/black, sqlfluff, SonarQube
Schema validation: Pydantic
Dependency & environment management: uv, Poetry
CI/CD & infrastructure: GitHub Actions, Docker, Kubernetes
AI-Assisted Development: Cursor, Claude Code, ChatGPT Enterprise
Future direction: agentic AI systems, LangChain, Azure OpenAI integration
What else you should know:
Team: Data Architecture Lead, Data Engineers, DataOps Engineers, Backend Engineer, Product Owner, collaboration with Frontend Engineers and Data Science and AI Engineers
Distributed team across Europe and India
Agile, collaborative environment; given the platform's organization-wide impact, we're looking for a mature, proactive, results-driven approach
Code quality is enforced through testing, typing, and tooling - not just code review
We work on multiple interesting projects at a time, so it may happen that we’ll invite you to an interview for another project if we see that your competencies and profile are well suited for it.
Your role
This is a results-driven contributor role, split roughly between hands-on engineering delivery and continuous improvement of the platform.
As a part of the project team, you will be responsible for:
Engineering & delivery (~70%)
Design, build, and maintain batch and streaming data pipelines in Python, including their orchestration, scheduling, and monitoring (Airflow)
Write reusable, well-typed Python libraries and internal packages used by other engineers and analysts
Build Python services and APIs (FastAPI) that expose data to downstream applications, and integrate with third-party and internal APIs
Develop data transformation logic in SQL and dbt on Snowflake, and build data models that support fast, reliable querying
Write unit, integration, and data-contract tests, and keep pipelines covered by automated CI
Profile and optimize Python code and data processing jobs for runtime, memory, and cost
Deploy and operate code in the cloud using containers, infrastructure as code, and CI/CD
Enforce access controls, secrets handling, and sensitive-data protections throughout the data lifecycle
Using AI coding assistants effectively while validating all generated output before it reaches production and helping build automated quality gates
Continuous improvement (~30%)
Find and fix efficiency, reliability, cost, and correctness issues in existing pipelines, refactoring toward simpler designs
Replace one-off scripts and notebooks with tested, packaged, scheduled code
Implement data quality checks, validation, and monitoring that catch issues before consumers do
Create matching logic to deduplicate and connect entities across multiple data sources
Document data processes and system architecture, and maintain project documentation
Do we have a match?
As a Data Engineer, you must meet the following criteria:
Strong Python experience: data structures, typing, error handling, generators/iterators, context managers, and the standard library
Software engineering fundamentals: modular design, dependency management, packaging, and API design
Testing discipline: pytest, fixtures, mocking, and writing code that's testable by construction
Hands-on experience building and operating ETL/ELT pipelines in production, not just scripts or notebooks
Strong experience with Snowflake and dbt
Experience with Apache Airflow or similar code-based orchestration tools
Solid working knowledge of SQL and data modeling, sufficient to design robust database schemas and query them effectively
Experience with Docker, Kubernetes, and CI/CD practices
Debugging and profiling skills; able to reason about performance, concurrency, and memory in Python
Experience with at least one public cloud (AWS or Azure)
Experience with version control systems (Git)
Able to explain technical trade-offs to both technical and business audiences
Experience using AI coding assistants such as Claude Code, Cursor, or similar on a daily basis
Strong communication skills and good knowledge of English (minimum C1 level)
Beyond the criteria above, we would appreciate the following nice-to-haves:
Experience with Apache Spark, ideally on Databricks
Experience with Pydantic or similar schema-validation libraries
Python web/API frameworks (FastAPI, Flask)
Async Python, multiprocessing, or other concurrency patterns
Experience with Azure AI Search or AWS OpenSearch
A second language: Go, Rust, Scala, or TypeScript
Familiarity with LLMs, Azure OpenAI, or agentic AI systems
More reasons to join us
Flexible working hours and approach to work: fully remotely, in the office or hybrid
Professional growth supported by internal training sessions and a training budget
Solid onboarding with a hands-on approach to give you an easy start
A great atmosphere among professionals who are passionate about their work
The ability to change the project you work on
- Department
- Observability Division
- Locations
- Poland
- Remote status
- Fully Remote