Role Overview
As a key member of the data team at Stylumia, you will develop and maintain web scraping scripts to extract data from various sources to support our Retail AI solutions. You will collaborate with cross-functional teams to implement effective scraping strategies and ensure data accuracy and reliability for downstream applications.
Responsibilities
- Develop and maintain web scraping scripts to extract data from a variety of sources.
- Collaborate with the data team to understand data requirements and implement effective scraping strategies.
- Conduct data quality assessments to ensure the accuracy and reliability of scraped data.
- Optimize scraping processes for efficiency and performance.
- Troubleshoot and resolve issues related to data extraction and scraping.
- Implement data storage and management solutions using SQL, Git, Redis, and Elasticsearch.
- Collaborate with cross-functional teams to integrate scraped data into downstream applications and systems.
Requirements
- Bachelor’s degree in Computer Science, Data Science, or a related field.
- Strong programming skills with proficiency in Python and Node.js.
- Solid understanding of web scraping techniques and experience with libraries such as BeautifulSoup, Scrapy, or Puppeteer.
- Proficiency in SQL and experience working with relational databases.
- Familiarity with version control systems, particularly Git.
- Knowledge of key-value stores, such as Redis, for caching and data storage.
- Experience with Elasticsearch or other search engines for data indexing.
Skills
- Python
- Node.js
- Web Scraping
- SQL
- Elasticsearch
Nice to Have
- Experience with distributed data scraping and parallel processing techniques.
- Knowledge of data cleaning and preprocessing techniques.
- Understanding of cloud platforms such as AWS, GCP, or Azure.
- Experience with containerization technologies like Docker.
- Familiarity with data visualization tools and libraries.