We provide IT Staff Augmentation Services!

Data Scientist Resume

0/5 (Submit Your Rating)

Dearborn, MI

SUMMARY

  • Solid exposure in Big Data, Data Analysis, Statistics, Forecasts, and Information Technology.
  • Experience in developing predictive models using R and Python (Supervised, Unsupervised, Deep learning, NLP, Network Analysis and Pyspark).
  • Deep knowledge on Spark, Hadoop, Data science and various Relational Databases (SQL Server, MYSQL and Oracle).
  • Hands on experience in working with Ecosystems like Hive, Pig, Sqoop, Map Reduce, Oozie, and Hue.
  • Experience in working with NoSQL databases such as HBase and Cassandra.
  • Expertise in data cleaning, data analysis and Data Visualization using Python (Pandas, Numpy, Matplotlib and Seaborn Packages).
  • Hands on experience in connecting Databases using Python (sqlalchemy package).
  • Hands on experience in Data Modelling: Dimensional and Relational.
  • Hands on experience and knowledge in developing Java scripting, VB scripting and Shell scripting.
  • Deep knowledge in creating custom objects in QLIKVIEW.
  • Deep Knowledge in Alteryx Tools and creating Macros in Alteryx.
  • Excellent team player with very good written and verbal communication skills.

TECHNICAL SKILLS

Analysis: R, Python, Tableau, Alteryx, Qlikview and MS - Excel

Databases: SQL Server, MYSQL, Oracle, Hbase and Cassandra

Languages: Python, Java, Scala, Pig Latin and Hive

Big Data Technologies: Hadoop, Spark (core, Spark, SQL, MLLIB and Pyspark), HDFS, Map Reduce, Hive, Pig, HBase, Sqoop, Oozie, Kafka and Flume

Operating system: Windows 8, Windows 7, UNIX, Linux and CentOS

Scripting Languages: UNIX Shell scripting, Java Script and VB script

Domain Knowledge: Streaming analytics, Big Data, Data science, IOT and Visualization

Tools: Git, MYSQL Workbench, Toad for MY SQL, SPSS Statistics, Minitab, Qlikview, Alteryx, Eclipse, IntelliJ IDE, Cloudera CDH5 and Horton Works.

PROFESSIONAL EXPERIENCE

Confidential, Dearborn, MI

Data Scientist

Responsibilities:

  • Built a classification prediction model in R to classify the Customs Data.
  • Conducted a Training session on how to analyze and process the data using Pig Latin and HiveQL (Hadoop Ecosystem).
  • Developed a pig script to blend data and show it in a Qlikview Dashboard.
  • Built a Data modelling (Dimensional and Relational modelling).
  • Developed a Macro in Alteryx using R (HTTR package). Macro will upload a file from Local Disk or Network Drive to SharePoint Libraries. Macro eliminates latency problem in Alteryx (Windows Authentication) to SharePoint
  • Developed Alteryx security tool (xml parse and R). Tool provides summary report about the tools used in workflows (lists, data/table/query and connection) and act as a firewall (check memory utilization, Load to the server, Personal data and access issue) to stop workflow before publish to server.
  • Built a custom object in Qlikview using JavaScript to get user input from the Dashboard and send it to Alteryx server to process the data and store it in a Database.
  • Automated the Alteryx software license File installation by developing a VBScript and T-SQL.
  • Conducted a Training session on how to build a custom object in Qlikview.
  • Proficient in R, JavaScript, VBScript, HDFS, Hive QL, Pig Latin, Sqoop and SQL.

Confidential, Bay Area, CA

Solutions Consultant

Responsibilities:

  • Worked on Flight prediction model, joining different datasets (flight data and weather data) using Spark SQL and Hive QL, processed a weather data in Scala.
  • Built a predictive model using Random forest algorithm in R to find Driver alertness in probability based on various factors Human, Vehicle and Environmental factors and published in the Dashboard.
  • Parsed a Json dataset and converted to csv file format using spark SQL and R
  • Performed network word count using Spark streaming
  • Performed a server log analysis using Spark and wrote the result to MySQL Database.
  • Build an IOT Model in Confidential software that will fetch data from SFTP server and push the data to HDFS target. Apache spark service will fetch the data from HDFS, will process the data using predictive model (Random forest and linear regression) based on key factor indicators and published the result to the dashboard. Dashboard will display the result based on population level, device level and group level.
  • Written a Scala code to fetch the current and Future weather data from the API to store it database. Weather data was collected to use against store data for demand forecast.
  • Written a JavaScript function that parse Xml data from a REST URL.
  • Written a JavaScript function to show the sample sales data in bipartite widget.
  • Created a form using HTML 5 and validate the form using JavaScript function.
  • Created SVG diagram like paths, point of interest and show details when Hover on Point of Interest in Geo-map overlay using Dojo JavaScript.

Confidential

Academic Internship

Responsibilities:

  • Proposed changes in project change management process of ERP Systems to improve business.
  • Tested the solution developed by the engineering team, with the user specifications through the Oracle Testing Module to ensure the services met user expectations.

We'd love your feedback!