Role Overview
Factori is a rapidly growing location and consumer intelligence company built for an AI-first world. We are looking for a Data Engineer with 3–6 years of experience to design, build, and operate large-scale data platforms. You will develop high-performance batch and streaming pipelines, optimize distributed data processing systems, and contribute to the infrastructure that powers our AI and analytics products.
Responsibilities
- Build and maintain scalable batch and streaming data pipelines using Apache Spark and distributed data technologies.
- Design, develop, and optimize ETL workflows with a focus on reliability, performance, and data quality.
- Develop efficient SQL queries, data models, and storage solutions across systems such as BigQuery, PostgreSQL, Elasticsearch, and DuckDB.
- Build and maintain workflow orchestration using Airflow.
- Monitor, troubleshoot, and optimize production data pipelines.
- Collaborate with product and engineering teams to design and deliver resilient, production-grade data systems.
- Leverage AI-assisted development tools to improve development speed and code quality.
Requirements
- 3–6 years of experience building distributed data processing systems or large-scale data platforms.
- Strong understanding of Apache Spark, distributed computing concepts, ETL design, HDFS, and SQL.
- Hands-on experience with relational and analytical databases such as PostgreSQL, Elasticsearch, or DuckDB.
- Proficiency in Python or Java with strong software engineering fundamentals.
- Experience with Airflow and at least one public cloud platform (GCP preferred).
- Familiarity with Linux environments, Bash scripting, and Git.
- Experience with data formats such as Parquet, Avro, or ORC is a plus.
- Hands-on experience using AI-assisted development tools like GitHub Copilot or ChatGPT.
Skills
- Apache Spark
- Python
- SQL
- Airflow
- GCP
Nice to Have
- Experience with Kafka, Pub/Sub, or other streaming technologies.
- Experience with Kubernetes, Docker, or infrastructure-as-code tools.
- Exposure to data warehousing technologies such as BigQuery, Snowflake, or Redshift.
- Understanding of CI/CD and observability.
Benefits
- Early impact – Help shape the tech stack and build products from the ground up.
- Agile culture – Small teams, zero bureaucracy.
- Great benefits – Group Health Insurance, daily breakfast, Friday team lunch, Fun O’Clock Fridays, and unlimited coffee, tea & snacks.