Recraft
ML Data Engineer
Salary
Competitive salary
Work type
Onsite
Level
mid
Category
Data Engineering
About the role
About Us
Founded in the US in 2022 and now based in London, UK, Recraft is an AI tool for professional designers, illustrators, and marketers, setting a new standard for excellence in image generation.
We designed a tool that lets creators quickly generate and iterate original images, vector art, illustrations, icons, and 3D graphics with AI. Over 3 million users across 200 countries have produced hundreds of millions of images using Recraft, and we’re just getting started.
Join a universe of professional opportunities, develop and support large-scale projects, and shape the future of creativity. We are committed to making Recraft an essential, daily tool for every designer and setting the industry standard. Our mission is to ensure that creators can fully control their creative process with AI, providing them with innovative tools to turn ideas into reality.
If you’re passionate about pushing the boundaries of AI, we want you on board!
Job Description
At Recraft, we’re building the next generation of generative models across images and text. We’re looking for an ML Data Engineer to scale our data pipelines for unstructured data (primarily images) and keep our training flows fast, reliable, and repeatable. You’ll design and operate high-throughput ingestion and preprocessing on Kubernetes, evolve our internal data-pipeline framework, and work hand-in-hand with ML engineers to ship datasets that move model quality forward.
Key Responsibilities
Develop and maintain data-ingestion pipelines to source and prepare large-scale image (and occasional text/HTML) datasets from open, publicly accessible, and permitted sources.
Own the end-to-end flow: raw data → quality/beauty/relevance filtering → dedup/validation → ready-to-train artifacts.
Operate and improve our Kubernetes-based data-pipeline framework (distributed jobs, retries, monitoring, automation).
Work with S3-style object storage: efficient layouts, lifecycle, throughput, and cost awareness.
Add tooling around pipelines (progress/health visualization, metrics, alerts) for observability and faster iteration.
Collaborate closely with ML engineers to align datasets with training needs and accelerate experimentation.
Requirements
Must-have
Strong Python fundamentals; you write clean, maintainable, production-ready code.
Solid hands-on Kubernetes experience (containers, jobs, batch/distributed processing).
Proven track record with unstructured data, especially images (loading, filtering, transforming at scale).
Experience developing data-ingestion or parsing tools for publicly accessible sources, including handling real-world reliability and failure cases gracefully.
Comfort with S3/object storage and moving lots of data efficiently and safely.
Pragmatic, detail-oriented, ownership mindset; you enjoy making systems reliable and fast.
Nice-to-have
Familiarity with ML workflows (PyTorch) and downstream training considerations.
Experience with image quality scoring, captioning, or image-to-text pipelines.
DAG/workflow visualizations or pipeline UX tooling.
DevOps fluency: Docker, CI/CD, infra automation.
What We Offer
Competitive salary and equity.
We’re able to offer Skilled Worker visa sponsorship in the UK for qualified candidates.
Real impact on model quality: your pipelines directly power training runs and product improvements.
Ownership with support: autonomy to design and improve systems, alongside experienced ML peers.
Modern stack: Python, Kubernetes, S3, internal pipeline framework built for scale.
Growth: a fast-moving environment where shipping well-engineered systems is the norm.
Tech stack
Motia
Data Scientist
Competitive salary
Leeds · 12h ago
You Trust
Business Analyst
Up to £46k
Fareham · 12h ago

Salutem Care and Education
Finance Business Intelligence Analyst
£40k – £50k
Windsor · 1d ago
Havebury Homes
Business Analyst
Up to £52k
Bury St Edmunds · 1d ago
Maldon Salt
Category Analyst
Competitive salary
Maldon · 4d ago
Bank of China
Business Analyst
Competitive salary
London · 4d ago
Golden Charter
Data Analyst
Up to £28k
Glasgow · 5d ago
Octavius
Junior Data Analyst
Competitive salary
Reigate · 5d ago

Qubitra
Senior Engineer
Competitive salary
Remote · 6d ago
ITH Systems
Senior AI Engineer
£75k – £85k
London · 6d ago
Thera Trust
Information Systems Analyst
Up to £38k
Grantham · 1w ago

Forge Holiday Group
Analytics Engineer
Up to £55k
Chester · 1w ago
Port of Dover
Data Placement Student
Up to £22k
Dover · 1w ago
Historic England
Marine Data Analyst (MDE Heritage Accelerator)
Up to £33k
Swindon · 1w ago
Echo
Business Analyst
£40k – £45k
Walsall · 1w ago
Condé Nast
Senior Data Analyst
Competitive salary
London · 3w ago
StarCompliance
Head of Data
Competitive salary
Remote · 3w ago
Brego
Lead Data Scientist
£90k – £110k
Remote · 3w ago
Compare the Market
Senior Product Analyst
Competitive salary
London · 3w ago
Polaris
Data Analyst
Up to £28k
Bromsgrove · 3w ago
Ground Control
Lead Data Engineer
Competitive salary
Billericay · 1mo ago
HelloFresh
Lead Data Analyst
Competitive salary
London · 1mo ago
Cooper Parry
Business Data Analyst
Competitive salary
Derby · 1mo ago
Medpace
Data Engineer
Competitive salary
London · 1mo ago
Recraft
ML Data Engineer
Salary
Competitive salary
Work type
Onsite
Level
mid
Category
Data Engineering
About the role
About Us
Founded in the US in 2022 and now based in London, UK, Recraft is an AI tool for professional designers, illustrators, and marketers, setting a new standard for excellence in image generation.
We designed a tool that lets creators quickly generate and iterate original images, vector art, illustrations, icons, and 3D graphics with AI. Over 3 million users across 200 countries have produced hundreds of millions of images using Recraft, and we’re just getting started.
Join a universe of professional opportunities, develop and support large-scale projects, and shape the future of creativity. We are committed to making Recraft an essential, daily tool for every designer and setting the industry standard. Our mission is to ensure that creators can fully control their creative process with AI, providing them with innovative tools to turn ideas into reality.
If you’re passionate about pushing the boundaries of AI, we want you on board!
Job Description
At Recraft, we’re building the next generation of generative models across images and text. We’re looking for an ML Data Engineer to scale our data pipelines for unstructured data (primarily images) and keep our training flows fast, reliable, and repeatable. You’ll design and operate high-throughput ingestion and preprocessing on Kubernetes, evolve our internal data-pipeline framework, and work hand-in-hand with ML engineers to ship datasets that move model quality forward.
Key Responsibilities
Develop and maintain data-ingestion pipelines to source and prepare large-scale image (and occasional text/HTML) datasets from open, publicly accessible, and permitted sources.
Own the end-to-end flow: raw data → quality/beauty/relevance filtering → dedup/validation → ready-to-train artifacts.
Operate and improve our Kubernetes-based data-pipeline framework (distributed jobs, retries, monitoring, automation).
Work with S3-style object storage: efficient layouts, lifecycle, throughput, and cost awareness.
Add tooling around pipelines (progress/health visualization, metrics, alerts) for observability and faster iteration.
Collaborate closely with ML engineers to align datasets with training needs and accelerate experimentation.
Requirements
Must-have
Strong Python fundamentals; you write clean, maintainable, production-ready code.
Solid hands-on Kubernetes experience (containers, jobs, batch/distributed processing).
Proven track record with unstructured data, especially images (loading, filtering, transforming at scale).
Experience developing data-ingestion or parsing tools for publicly accessible sources, including handling real-world reliability and failure cases gracefully.
Comfort with S3/object storage and moving lots of data efficiently and safely.
Pragmatic, detail-oriented, ownership mindset; you enjoy making systems reliable and fast.
Nice-to-have
Familiarity with ML workflows (PyTorch) and downstream training considerations.
Experience with image quality scoring, captioning, or image-to-text pipelines.
DAG/workflow visualizations or pipeline UX tooling.
DevOps fluency: Docker, CI/CD, infra automation.
What We Offer
Competitive salary and equity.
We’re able to offer Skilled Worker visa sponsorship in the UK for qualified candidates.
Real impact on model quality: your pipelines directly power training runs and product improvements.
Ownership with support: autonomy to design and improve systems, alongside experienced ML peers.
Modern stack: Python, Kubernetes, S3, internal pipeline framework built for scale.
Growth: a fast-moving environment where shipping well-engineered systems is the norm.