We provide IT Staff Augmentation Services!

Data Science & Analytics Consultant (sr.) Resume

0/5 (Submit Your Rating)

Des Moines, IA

SUMMARY

  • Data Scientist with 4+ years of experience in implementing machine learning models.
  • Overall 8 years of experience in the IT field that includes implementing data pipelines and warehousing solutions.
  • SME implementing on end - to-end Machine Learning projects from project discovery to deployment/production.
  • Expert in developing code for both REPL phase (exploration) of Data Science projects as well as batch-submission of packages during deployment.
  • Proficient in Multivariate Regression, Principal Component Analysis, Decision Trees and Random Forests, Boosted Trees, Clustering.
  • Experience building models in scenarios that have persisting issues like non-linearity and class-imbalance.
  • Worked in diverse industry domains like Banking, Home Mortgage, Insurance, Retail and Manufacturing.
  • Understands business problems and can formulate them into well-posed Machine Learning problems in order to provide actionable insights.
  • 3+ years’ experience with Python and a variety of Machine Learning Python, H2O and Apache Spark libraries like scikit-learn and Mllib.
  • 8 years experience in writing SQL queries and building SQL tables and databases.
  • Experienced in data scrubbing, data quality maintenance, data integration, data mapping, data profiling and data validation.
  • Collaborated with teams that aimed to build distributed, scalable data pipelines to ingest and process data.
  • Skilled in Scala and Python and familiarity with Hadoop framework and Apache Spark.
  • Extensive knowledge in programming with Resilient Distributed Datasets (RDDs) as well as DataFrames.
  • Developed workflows using Hive and Sqoop.
  • Good Knowledge on Amazon EC2, EMR, S3, RedShift.
  • Expertise in using Flume in collecting, aggregating and loading log data from multiple sources into HDFS.
  • Developed Spark applications using scala and Spark-SQL/Streaming for faster testing and data processing.
  • Collaborated effectively with other data scientists and analysts across the organization and ensured accurate and timely fulfillment of deliverables.
  • Experience working on cloud platforms namely cloudera, Azure and Hortonworks.
  • Experience in development and deployment of predictive models in cloud or distributed systems.
  • Strong experience and knowledge in Data Visualization with Tableau creating Line and scatter plots, Bar Charts, Histograms, Pie chart, Dot charts, Box plots, Time series, Error Bars, Multiple Charts types, Multiple Axes, subplots etc.
  • Experience working in Agile and Waterfall environments.
  • Possess excellent organization skills, documentation skills and technical report writing skills that were proven to be highly beneficial in highly regulated sectors like Banking.

TECHNICAL SKILLS

Programming Languages: Scala, Python, Java, SQL, SAS, Hive-QL

Big Data: SQOOP, Hive, Apache Spark, H2O, Kafka, Flume

Frameworks: Hadoop HDFS2.0, Cloudera, Hortonworks

ML Algorithms: Linear Regression, Logistic Regression, Cox Regression, XGBoost, FP-Growth, Nerual Networks, Feature Engineering, Dimensionality Reduction

Databases: HBase, Hive, SQLServer, Teradata, DB2, Aster, SQL Server 2008, MySQL

Distributed File Systems: HDFS, S3

Tools: and Software: Tableau, SAS E-Miner, IBM SPSS, SAS Visual Analytics, Teradata Studio (Aster), Power BI, Alteryx, QlikView

Open Source Libraries: Numpy, pandas, scikitlearn, scipy, hmmlearn, Spark Mllib for regression, classification, frequent pattern mining, model selection, H2O (like Word2Vec), sparkling H2O

ETL Tools: Informatica Power Centre, Data, AbInitio, Talend, MS SQL Analysis Manager, DB2 OLAP, CloverETL, QlikView

Cluster Managers: YARN, Mesos

IDEs and Platforms: PyCharm, IntelliJ, Spyder, Jupyter Notebooks, Eclipse

SAS Packages: SAS GRAPH, SAS/ACCESS, SAS IML, SAS/STATS, SAS ODS, BASE SAS

PROFESSIONAL EXPERIENCE

Data Science & Analytics Consultant (Sr.)

Confidential, Des Moines, IA

Responsibilities:

  • Developed code for merging datasets, data preparation, reporting, and deployment of models.
  • Responsible for deploying and maintaining a machine learning model that helps predict an incoming complaint from home loan customer.
  • Frequently meet with the business teams to gather inputs on a problem and develop model that caters to the solution.
  • Achieved completion of an MBO for 2018 by reduction of complaints and improvement in customer experience through machine learning.
  • Go-to team member for open source data ecosystem maintenance for the team.
  • Worked on many modeling projects aimed at predicting customer complaints using machine learnig algorithms like Logistic Regression, Cox Regression, PCA (for dimensionality reduction) and Markov Chain (AI POC).
  • Additionally, used another open source analytical engine H2O or development of models.
  • Occasionally used SAS for developing time-dependent covariates’ model (Ex: Cox Regression).
  • Loading data from different datasets and deciding on which file format is efficient for a task.
  • Troubleshoot and debug any Hadoop ecosystem run time issues.
  • Worked on creating Spark jobs that process the true source files and successful in performing various transformations on the source data using Spark Dataframe/Dataset, Spark SQL API’s.
  • Collaborated with the engineering team in creating Hive tables that better enhance the performance of various queries run on these Hive tables.
  • Involved in the establishing connections between Apache Spark and MySQL using JDBC connectors and establishing a mapping between HDFS files and MySQL and Teradata tables.
  • Extensively worked with Avro and Parquet files and converted the data from either format parsed semi-structured JSON data and converted to Parquet using Dataframes in Spark.
  • Exploring with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Spark on YARN.
  • Adhered to data regulations and user privacy rules while working on data with sensitive customer information.

Environment: HDFS, Hive, Kafka, Sqoop, Python, Unix, MySQL, scala, python, pyspark, Teradata SQL, SAS, H2O

Data Engineer

Confidential, Chicago, IL

Responsibilities:

  • Part of the Big Data team in generating data for the reports and designing data workflows.
  • Involved since the early stages of the creation of a data lake.
  • Worked on the development of Data Warehouse, Business Intelligence architecture that involves data integration and the conversion of data from multiple sources and platforms.
  • Developed models that enabled the pricing team to reduce the gap between the estimated cost and the actual cost.
  • Involved in the process of load, transform and analyze data from various sources into HDFS (Hadoop Distributed File System) using Hive and Sqoop.
  • Worked with Linux systems and RDBMS database on a regular basis in order to ingest data using Sqoop.
  • Managed a customer base worth a net value of $100 million for their data analysis and reporting purposes using Python and Tableau as per the C level executives’ requirements.
  • Extracted data from different plants’ ERP systems and manipulated as per the end requirement. Documented the metadata details for future reference.
  • Developed data pipelines for operations, production, sales, finance and marketing.
  • Worked extensively on data migration during the organization expansion in Central America by collecting data from the new companies.
  • Involved in performance tuning of the ETL process by addressing various performance issues at the extraction and transformation stages.
  • Designed and implemented HIVE queries and functions for evaluation, filtering, loading and storing of data.
  • Created dashboards on Tableau extracting data from different sources using data blending from Axapta, SAP, Oracle, SQL Server, MS Access and CSV at single instance.
  • Involved in loading data from UNIX file system and FTP to HDFS.
  • Worked in spark ecosystem and used Spark SQL functionality extensively and developed scala programs for data extraction and transfer on different file formats like Text file, CSV file.
  • Developed Spark code using Scala and Spark-SQL for faster testing and data processing.

Environment: HDFS, Hive, Sqoop, Python, Unix, SQL Talend, Tableau, Scala, SQL Server, AbInitio, Alteryx, Python Scikit-learn.

ETL Developer

Confidential

Responsibilities:

  • Cleaned, analyzed and selected data to gauge customer experience.
  • Implemented new statistical and mathematical methodologies as needed for specific models and analysis.
  • Used algorithms and programming to efficiently go through large datasets and apply treatments, filters, and conditions as needed.
  • Created meaningful data visualizations to communicate findings and relate them back to how they create business impact.
  • Involved in data modeling, data architecting and ETL in the corporate environment and implemented modular and reusable codes to run across 20+ product groups.
  • Updated all vendor information,retailprices, coupon and item descriptions.
  • Created exception based reports to pinpoint errors in company master data.
  • Wrote SQL queries to gather point of purchase and sales information.
  • Composed interactive report, which provided regional sales managers a snapshot of progress in a particular region.
  • Utilized ODBC for connectivity to databases and MS Excel for automating reports and graphical representation of data to the Business and Operational Analysts.
  • Calculated on time delivery percentages for top vendors.
  • Trained merchandising assistants in usage and creation of time based sales reports for merchandising department.
  • Measured effectiveness of various marketing methods in growth of company to take informed decisions on marketing strategies.
  • Analyzed data of customers at various stages like enquiry and booking stage to rate customers.

Environment: Informatica Power Center 8.6/9.0.1, Oracle 10g, UNIX, MS SQL Server 2008, MS-Access, Autosys

ETL/BI Developer

Confidential

Responsibilities:

  • Worked for various clients in providing ETL and data warehousing support.
  • Extensively used Star Schema methodologies in building and designing the logical data model into Dimensional Models.
  • Involved with Data Analysis primarily identifying data sets, source data, source metadata, data definitions and data formats.
  • Developed Informatica mappings, sessions, workflows and have written PL SQL codes for effective and optimized data flow.
  • Conducted GAP analysis in the identification ofbusinessrules,businesssystem process flows, creating AS-IS/TO-BE mapping for all new enhancements.
  • Created requirements artifacts like theBusinessRequirements Document (BRD) and Functional Requirements Document (FRD).
  • Assisted the project manager in determining the scope and Work Breakdown Structure (WBS) to meet project milestones.
  • Responsible to tune ETL procedures and STAR schemas to optimize load and query performance.
  • Constantly monitoring the application by performing SQL queries with the database and managed the data Performed data analysis using SQL queries and used reports to understand current data metrics and provided information tobusinessusers.
  • Developed Data Mapping, Data Governance, and Transformation and cleansing rules for the Master Data Management Architecture involving OLTP, ODS.
  • DataStage jobs were scheduled, monitored, performance of individual stages was analyzed and multiple instances of a job were run using DataStage Director.
  • Developed Informatica mappings, sessions, workflows and have written Pl SQL codes for effective and optimized data flow coding.
  • Responsible for generating reports through Tableau and facilitated in the decision making process of the senior management including dashboards and summary reports.
  • Developed and maintained end user documents for applications and trained stakeholders of project.
  • Designed, developed and implemented test plans, test scenarios and test cases.
  • Performed analysis to measure clients’ performance at various stages and enabled data driven decisions.

Environment: Informatica (Power Center Repository Manager, Designer, Workflow Manager, and Workflow Monitor), Teradata, Oracle 11g, SQL, IBM DB2, UNIX, Citrix, Putty, TOAD.

We'd love your feedback!