Role Overview
At Accendra Health, we understand that healthcare is complex, and we’re here to make it easier. We are seeking a highly skilled Data Scientist with a strong background in machine learning, data engineering, and model optimization. The ideal candidate should be proficient in Python, PySpark, and SQL, experienced in time series forecasting, feature engineering, and data model performance evaluation, and capable of working with large-scale data integration projects across various domains.
Responsibilities
- Develop and optimize machine learning models with a focus on time series forecasting and predictive analytics.
- Perform feature engineering and data model optimization to enhance model accuracy and efficiency.
- Continuously evaluate model performance using metrics such as MAPE, RMSE, R², and adjust strategies accordingly.
- Build and implement data pipelines using PySpark, SQL, and cloud-based solutions for seamless data integration.
- Work on large-scale data integration projects, leveraging tools such as Boomi, SnapLogic, SSIS, or Palantir to extract, transform, and load data.
- Utilize Palantir Foundry, Google Cloud, AutoAI, and Google Colab for data modeling, processing, and automation.
- Design and maintain data warehouse solutions to support advanced analytics and business intelligence.
- Perform complex data transformations using SQL queries and data objects to support AI/ML-driven initiatives.
- Collaborate closely with business stakeholders to ensure models align with user expectations and business objectives.
- Deploy, monitor, and continuously improve machine learning models in production environments.
- Communicate technical findings and insights effectively to both technical and non-technical audiences.
Requirements
- Proficiency in Python, PySpark, and SQL for data analysis, feature engineering, and model development.
- Expertise in time series forecasting models, including ARIMA, Prophet, LSTMs, and ML-based approaches.
- Strong experience in data model optimization, feature engineering, and performance evaluation.
- Deep understanding of ML model evaluation metrics and best practices in improving model accuracy.
- Hands-on experience in data engineering, working on data pipelines, ETL, and data transformation projects.
- Experience using Boomi, SnapLogic, SSIS, or Palantir for data integration.
- Proficiency in cloud computing, particularly Google Cloud (BigQuery, Vertex AI, Cloud Functions, etc.).
- Experience with Palantir Foundry for data processing, analysis, and visualization.
- Ability to optimize and query large-scale datasets using data lakes and relational databases.
- Familiarity with AutoAI for automated model selection and hyperparameter tuning.
- Experience with Google Colab for collaborative machine learning development.
- Excellent problem-solving and communication skills.
Skills
- Python
- PySpark
- SQL
- Google Cloud
- Machine Learning