Job Description
About Us
STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains. We design, develop and manufacture high-performance unmanned systems that are software-defined, mass-scalable, and cost-effective. This provides our operators with a decisive edge in highly contested environments.
We're focused on delivering deployable, high-performance systems — not future promises. In a time of rising threats, STARK is bolstering the technological edge of NATO Allies and their Partners to deter aggression and defend Europe — today.
About the team
The Data Operations team owns the entire data lifecycle behind STARK's AI stack: collection, acquisition, generation, curation, and management. We run our own data-collection campaigns across Europe, evaluate new sensors and platforms, and build the internal data platform that turns raw recordings into ready-to-use datasets. Everything we produce feeds directly into the perception and autonomy systems deployed on STARK's platforms — a real data advantage is built, not bought. The team is scaling up right now: real scope, direct impact, no legacy.
Your mission
Data is the fuel of STARK’s AI stack — you build the engine that makes it usable. You own the software backbone of our data platform: the metadata systems, ETL pipelines, data contracts, catalogs, databases, and internal tools that let engineers find, understand, validate, and reuse terabytes of multi-sensor field data in minutes, not days. You treat data context as a product: structured, searchable, version-aware, documented, and traceable from raw recording to processed asset, annotation delivery, dataset, and downstream ML workflow. Today, much of this is manual, scattered, or implicit — your job is to automate it away, support labeling efforts with the right data tooling, and turn operational data into reliable systems.
Responsibilities
-
Design, implement, and maintain our metadata database and data catalog (datasets, recordings, sensors, labels, lineage)
-
Build and operate ETL/ingest pipelines that bring field recordings, synthetic data, and external deliveries into our cloud storage (GCP)
-
Own the data management and labeling lifecycle end-to-end: coordinate and communicate with external labeling companies and data subcontractors, track deliveries, run QA reports, and build the operational workflows they work in
-
Develop internal enabling tools for the whole AI organization: dataset search and filtering, APIs/backend, dashboards, and self-service data access
-
Run data migrations and indexing jobs; keep the catalog consistent and fast as data volume grows
-
Handle admin support and user access management — and then automate these support tasks so they stop being manual work
-
Establish good engineering hygiene in a young codebase: tests, typing, docs, logging, CI/CD
-
Shape the long-term architecture and vision of the data platform together with the team
Qualifications
-
Strong Python
-
Solid SQL/PostgreSQL, including schema design
-
Experience with data modeling and metadata systems
-
Experience designing and operating ETL/data pipelines
-
Docker and CI/CD basics
-
Hands-on with object storage (GCS, S3, or similar)
-
Good software engineering hygiene: tests, docs, typing, logging
-
Organized and pragmatic: you can prioritize between a quick fix and a proper solution, and you know when each is right
-
Not allergic to support tasks — but technical enough to automate the support away
-
Comfortable coordinating with external vendors and non-technical stakeholders
Nice to have
-
Familiarity with ML datasets and labeling workflows (images, video, lidar; annotation formats like COCO)
-
Experience with synthetic data generation or GenAI-assisted data workflows (auto-labeling, data augmentation, foundation-model-based curation)
-
Experience with GCP services beyond storage (BigQuery, Cloud Run, IAM)
-
Experience with data versioning / dataset tooling (DVC, LakeFS, FiftyOne, or similar)
-
Experience in a startup environment — comfortable with ambiguity and changing priorities
-
Exposure to robotics data formats (ROS bags, MCAP, PX4 logs)
Skills & Technologies
Other Jobs at STARK Defence
Similar Opportunities
GNC/Autopilot Engineering Manager (All genders)
STARK Defence
Flight Test Engineer (m/f/d)
STARK Defence
Data Engineering Intern (All genders)
STARK Defence
DMU (Digital Mock-Up) Engineer
Helsing
Senior Software Engineer - Go (All Genders)
STARK Defence
Production Data Engineer (all genders)
STARK Defence