We provide IT Staff Augmentation Services!

Sr Data Scientist | Applied Machine Learning/ai Architect Resume

4.00/5 (Submit Your Rating)

Chicago, IL

SUMMARY:

  • Have 17+ years of experience in architecting, designing and implementing large - scale systems and 8+ years’ experience serving as the lead architect and expert designer of Applied Machine Learning, AI/Deep Learning, Big data solutions and cloud-based solutions.
  • Designed and implemented scalable Big Data architecture solutions for various applications needs and worked with rapidly evolving technologies at different organizations to analyze and define unique solutions.
  • Have strong hands-on experience implementing ML, AI/Deep Learning, NLP and big data solutions with technologies including Python, R, Scala, TensorFlow, Karas, Caffe2, Scikit-learn, PyTorch, OpenCV, Spark MLlib, Spark Streaming, Kafka/Confluent, Spark SQL, H adoop, MapReduce, Sqoop, Flume, Pig, Hive, Impala, HDFS, NOSQL(MongoDB, HBase and Cassandra), AWS(EMR, Kinesis, DynamoDB, S3, Redshift, and Sagemaker), and Oozie.
  • Strong experience architecting solutions to migrate data from traditional EDW (Teradata & Oracle Exadata) to low-cost Hadoop EDW.

TECHNICAL SKILLS/PROFILE:

Python, Scala, R, Tensorflow, Keras, Scikit-learn, Mahout, Pandas, NLTK, Kafka, Storm, Sqoop, Flume, Spark, Spark streaming, Hadoop, MapReduce, Pig, Hive, HDFS, Hcatalog, Avro, Impala, YARN, Zookeeper, HBase, Cassandra, MangoDB, Oozie, R, RMR, Solr, Elastic Search, Neo4j, CDH, HDP, MapR, J2EE, Spring boot, JDBC, JMS, JNDI, SOAP, JAX-RPC, JAXP, XML, XSL, XSLT, HTML, AJAX, Log4J and JavaScript, Java Mail, JTS, JAAS, JTA, JNI

PROFESSIONAL EXPERIENCE:

Confidential, Chicago, IL

Sr Data Scientist | Applied Machine Learning/AI Architect

Responsibilities:

  • Architected, designed, implemented and deployed ML and Deep Learning systems to proactively monitor, predict application performance issues, to automate issue resolution and to perform predictive maintenance of the network devices and servers.
  • Spearheaded effort to collect, analyze, transform and feature engineer the data points for different ML/AI algorithms in big data and NOSQL environment.
  • Designed, evaluated and implemented solutions for handling high-volume real-time data streams using Kafka and Spark Streaming for real-time predictions.
  • Designed the architecture and performed big data analytics and data science experiments on AWS cloud using technologies S3, Kinesis, EMR, Redshift, DynamoDB, Sagemaker, Tensorflow, EC2 and Lambda.
  • Designed, implemented and deployed Deep Learning solution to predict critical network devices and servers’ health issues using RNN/LSTM algorithm and TensorFlow/Keras.
  • Designed, implemented and deployed Deep Learning solution to predict user’s sentiment (NLP) with RNN algorithm and Embedding using TensorFlow/Keras.
  • Designed and implemented solution to predict device failures by timelines using ML classification (Naive Bayes, Logistic Regression, Random Decision Forests ).
  • Predicted timeline for maintenance and failure using regression (SVM, Random Forests).
  • Performed feature engineering using different techniques and best practices (Including Interpolation, Adaptive binning, Box-cox transformation).
  • Fine-tuned hyper prams of experimented and production models for the best possible speed of convergence and scoring accuracy.
  • Worked closely with infrastructure team to perform capacity planning for Big data, AI and ML systems in production and test environments to help ML team to implement appropriate ML algorithms, run / train ML tests and experiments.
  • The developed solutions helped Confidential to prevent $50 million revenue loss that could have been caused by the downtime and application performance issues.
  • Created a vision for and provided technical guidance to the ML team to solve challenging problems in creating custom Artificial Intelligence for IT Operations (AIOps) system.
  • Worked closely with other domain experts to refine and execute Confidential ’s technical strategy in AI and ML.
  • Collaborated with researchers to combine domain knowledge with state-of-the-art deep learning techniques.
  • Advised internal leaders on recent deep learning advancements in the industry and academia to further influence research direction and business decisions.
  • Worked across diverse teams, perspectives and opinions and quickly built consensus on ML and AI solutions.
  • Experimented with different ML and AI approaches, models and solution options and built the custom solutions that best fit Confidential ’s AIOps needs.
  • Encouraged informed risk-taking and acted as a catalyst for innovation at Confidential and generated practical, sustainable and creative options to solve problems while maximizing existing resources.
  • Proactively developed, maintained and evangelized technical knowledge in Machine Learning and AI technologies. Implemented solutions following current trends and best practices.

Confidential, Lincolnshire, IL

Applied Data Scientist | Lead Big Data Architect

Responsibilities:

  • Responsible for leading efforts of modernizing and migrating legacy Analytics System to big data-based Analytics System.
  • Created solution, physical, technical and data architectures for Analytics Platform that provides users with reporting and ad-hoc querying capabilities on actual spend and net demand forecast against product parts and finished goods for Confidential suppliers and contract manufacturers.
  • Built architecture and technical solution to store and analyze highly complex and hierarchical products and parts data, aggregates and views using NoSQL database MongoDB .
  • Advised management team on tactical goals and also provided the long-term road map for IT and business systems and processes.
  • Provided direction and guidance to Confidential IT team on Big Data solution, tools and best practices.
  • Created strategic solution design that leverages current Confidential infrastructure and software components, minimizes TCO, decreases speed-to-market and drastically improves system performance and provides scalable EDW platform with Spark, Datameer, Hadoop, Hive and OBIEE technology stack.
  • Implemented critical solution components using technologies including Spark Streaming, Spark SQL, MongoDB, Datameer, Hadoop, MapReduce, Hive, Impala, HDFS, Sqoop, Oozie, Shell scripting and other big data technologies.
  • Designed and implemented ETL data pipelines on huge data sets using Hive and MapReduce.
  • Lead designing of infrastructure architecture and performed capacity planning to ensure Hadoop and Datameer clusters are strategically positioned to meet business goals, product vision, delivery time-lines and budget limits.
  • Designed and implemented data acquisition and ETL pipeline using Sqoop, Datameer, Spark, Spark Streaming and Hive technologies.
  • Selected right tools for data governance, data quality management and data lineage and mentored client’s IT team on usage of the tools.
  • Designed and implemented security and privacy strategy for data in HDFS. Created data compression and data compaction best practices and techniques.
  • Architected, designed and implemented data lake on AWS cloud platform with EMR, Data pipeline, RDS, Redshift, S3, VPC and EC2 services. Automated data governance and security with ElasticSearch, Dynamo DB, Lambda, Cognito, IAM, STS and KMS.

Confidential, Chicago, IL

Lead Big Data Architect

Responsibilities:

  • Responsible for leading the Big Data initiative at Confidential .
  • Created solutions that meets both tactical and strategic business goals.
  • Created solution and technical architecture for EDW and ETL data pipelines.
  • Designed and developed technical solution to store and analyze huge internal and external data sets with 100% availability using Cassandra NoSQL database;
  • Architected solution to migrate data from Informatica EDW to Hadoop EDW.
  • Slashed EDW cost by 40% by moving complex ETL flows to Hadoop cluster, while ensuring risks related to data quality and data completeness are mitigated.
  • Implemented critical solution components using technologies including Hadoop, MapReduce, Hive, Impala, HDFS, HBase, Sqoop and Flume.
  • Designed and implemented ETL data pipelines on huge data sets using Spark, Hive and MapReduce.
  • Involved in designing of infrastructure architecture and performed capacity planning to ensure Hadoop and NoSQL clusters are strategically positioned to meet business goals, vision, product roadmap, delivery time-lines and budget needs.
  • Implemented data acquisition using Sqoop and Flume technologies.
  • Designed and implemented security and privacy strategy for data in HDFS. Created data compression and data compaction best practices and techniques.

Confidential, Hoffman Estates, IL

Lead Big Data Architect

Responsibilities:

  • Created solution architecture, technology architecture and implemented solutions for different big data and machine learning initiatives. Implemented critical application components using technologies including Pig, Impala, Hive, MapReduce, HDFS, Cassandra, MongoDB, HBase, Spark, Spark Scala API, Spark streaming, Storm, Kafka, Flume, Sqoop, Twitter API, Facebook API, Elastic Search, R, rmr, Python, Tableau, Neo4j.
  • Architected solutions and created tactical and strategic roadmap to migrate data from Teradata to Hadoop EDW. This migration resulted in total cost saving of 40%.
  • Designed and implemented real-time predictive analytics and recommendation engine to solve business needs such as “merchandise prediction”, “item price optimization” and “item recommendations” using big data and machine learning technologies including Python, NumPy, SciPy and Pandas, R, Tableu, H2o, Spark MLlib.
  • Involved in entire lifecycle of machine learning projects including data preparation, predictive model creation, test and validate the model, deployment of the model in prod environment, analysis, monitoring and model optimization.
  • Strong hands-on experience with different machine learning algorithms, model fitting and model optimization techniques including clustering, classification, regression, ensemble learning, reinforce learning, time-series analysis, pattern recognition and cross validation.
  • Designed and implemented ETL flows on Hadoop using Pig, MR and Hive.
  • Implemented big data pipeline from external data sources using Sqoop and Flume technologies.
  • Developed technical architecture, designed and implemented enterprise wide centralized logging system using Elasticsearch, Logstash and Kibana.
  • Designed Cassandra, MongoDB data models for different business needs. Lead capacity planning effort for Cassandra and MongoDB clusters. Put together best practices for NoSQL key design, development and production deployment.
  • Implemented MongoDB based applications using Java Driver API, Casbah (Scala API) and Mongo shell commands.
  • Implemented Cassandra applications using Datastax Java API, Hector API and CQL API.
  • Built enterprise data hub following lambda architecture principles with a combination of technologies Hadoop, Spark and Spark streaming.
  • Implemented Spark and Spark streaming applications using Scala, Spark SQL and MLlib APIs.
  • Created reference architecture for different Big Data business case scenarios and put together development best practices and laid out a plan for data governance using Hadoop.
  • Designed and implemented security model for Hadoop cluster.
  • Implemented major components as a proof of architecture design and to provide direction to the application teams.
  • Created solution and technology architecture design strategy for data science and machine learning projects on big data platform to build critical enterprise wide integrated products and services.
  • Lead the architecture strategy and designing of critical initiative to get a real-time 360 degree view of the customers by integrating discrete and disparate data source while absorbing the data quality complexities from silo-ed data sources and systems. Built predictive models on integrated customer profiles to target the right customers with right products, service at the right time. And also engage the customers in the most effective manner.
  • Involved in infrastructure architecture and capacity planning to ensure infrastructure and Hadoop, Spark and NoSQL clusters align with over all enterprise architecture. Provided necessary recommendations to operations team.
  • Created strategy, designed solutions and POCs for service oriented big data architecture, big data security, real-time analytics and machine learning on big data.
  • Created best practices, techniques and tools to simplify and speed up new applications on big data.
  • Identified and evaluated new big data technologies/products/tools that help fill the gap in over all enterprise architecture for future business needs.

Confidential, Chicago, IL

Big Data Enterprise Architect

Responsibilities:

  • Designed and implemented machine learning projects including “Provider’s Claims Fraud Detection” and “Member Health Risk Prediction” using bigdata and machine learning technologies including Python, NumPy, SciPy and Pandas, R and Tableu.
  • Designed the solution architecture of this big data system that provides real-time predictive analysis based recommendations to the applicants and members using Hadoop technology stack.
  • Designed Cassandra data models and involved cluster capacity planning.
  • Implemented Cassandra clients using Java API and performed data bulk loading into Cassandra data store.
  • Created technical architecture for predictive analysis based recommendations using IBM Infosphere Big Insights, streams, SPSS modeler, CEP Engine, Time series module.
  • Designed architecture for alternative solution option using R, Hadoop and Python scientific libraries.
  • Massive amount of structured, unstructured and semi-structured data is ingested from member/applicant call logs, emails, portal click streams, health care research and clinical trial projects, social media, discussion boards as input for the predictive analysis.
  • Created prototypes for the major components as a proof of architecture design and to provide direction to the IT team.
  • Created solution architecture artifacts and deliverable using Sparx Enterprise Architect tool and presented the architecture to business and IT teams.
  • Involved in infrastructure and capacity planning to ensure the necessary infrastructure exists to support the solutions. Provided necessary recommendations.
  • Worked closely with data scientists, business users and marketing team for architecture realization.
  • Guided application development teams on implementation of big data components and modules.
  • Aligned solution architecture design with enterprise architecture context and strategic direction.
  • Worked closely with enterprise architecture governance, data architects and enterprise security to re-align data architecture for big data strategy.

Confidential, Hoffman Estates, IL

Big Data Solution Architect

Responsibilities:

  • Designed the solution architecture and lead the team to quickly deliver near real-time context and location aware item price recommendation engine that processes massive amount of enterprise wide data in Hadoop Cluster to target improved sales and the margins and build new customer base and to targets existing customer retention.
  • Created prototypes leading to specific engineering initiatives and lead the execution of technology strategy for decisions on technology platforms and partnerships.
  • Created technical solution that solves business need for a business rules engine within a business process flow that works against big data in Hadoop.
  • Made presentation on solution architecture for other technical teams and top level executives to provide guidance for their system architecture.
  • Modeled system architecture and design artifacts using UML2 modeling tool.
  • Setup standards and processes for Hadoop based application design and implementation.
  • Designed and implemented a metadata automation framework for big data in Hadoop.
  • Designed and implemented real-time big data analytics and reporting system using HBase and JasperSoft.
  • Created prototypes for alternative BI solutions using Microstrategy, Informatica and Pentaho-kettle ETL.
  • Designed and implemented reusable Hadoop workflow using Oozie that interacts with different Hadoop technologies.
  • Designed and implemented prediction model using Mahout Implementations of Collaborative filtering, clustering and classification algorithms.
  • Mentored different teams working on Hadoop and provided a direction for their implementation.
  • Drove the development of agile mapreduce programing practices.
  • Architected data acquisition and transformation solutions for Hadoop using Sqoop, Flume, custom and Informatica HParser.
  • Provided strategy and direction for unit testing and QA process with a focus on Hadoop and HBase components.
  • Put together a strategy and recommended technologies and for future business needs.
  • Setup the build and deployment infrastructure for the system using Maven2 and Nexus.
  • Created environment for continuous integration and code coverage using Hudson and Cobertura.
  • Designed and lead implementation of Drool rules engine and Hadoop integration to implement configurable pricing logic against big data.
  • Orchestrated business process engine using JBPM and integrated it with rules engine.
  • Setup version control infrastructure using Subversion (SVN) tool.

We'd love your feedback!