Data Analyst Resume
Hicksville, NY
PROFESSIONAL SUMMARY:
- Extensive 4+ years of experience in Data Modelling, Analysis, Migration and ETL with specialization in Python, R studio, Tableau, SQL, SPSS, Scala and Big Data frameworks like Spark, Spark - Sql, Hive, Pig, Sqoop, Kafka.
- Expertise in Hadoop eco system components HDFS, Ambari,Hive, Sqoop, Oozie, HBase and Pig for scalability, distributed and high-performance computing.
- Expertise in writing the complex Hive and Impala queries.
- Experience of converting HQL/SQL queries into Spark and Pig transformations using Spark RDDs and Pig Relations.
- Developed Spark code and Spark-SQL for faster testing and processing of data using PySpark and Scala.
- Hands-on experience of importing data from traditional databases in to HDFS and Hive using Sqoop and performing transformation using pig and spark for analysis and exporting results to tableau for visualizations.
- Experience in working with both Streaming and Batch data processing with the help of Spark.
- Experience in various machine learning algorithms like K-Means Clustering, Principal component analysis, XGBoost, Linear learner classification and regression.
- Experience in using libraries such asPandas, Numpy, Scikit-Learn and Matplotlib in python, Spark-SQL, Mlib, GraphX in spark.
- Experience in fetching and Monitoring data to identify KPI’s. Developed dashboards for based on KPI data to present data insights.
- Experience in translating data insights into business objectives and decisions based on statistical analysis.
- Well versed in various AWS application services like Recognition, Transcribe, Translate, Polly, Comprehend and Lex.
- Expertise with Amazon ML Platform SageMaker and ML frameworks like Pytorch, TensorFlow and Keras.
- Experience in web scarping using beautifulsoup in python for data collection and exploration.
TECHNICAL SKILLS:
Programminglanguages: Scala, SQL, Python
Databases: Oracle, MS SQL, PL/SQL,My SQL, HBase, MongoDB, Cassandra, Redshift, DynamoDB
Algorithms: Tableau, Python and R Studio HDFS, Sqoop, Hive, Pig, Oozie, Flume, Spark, Spark-sql Recognition, Transcribe, Translate, Polly, Comprehend and Lex, SageMaker, S3 Means Clustering, Principal component analysis, XGBoost, Linear learner classification and regression.
PROFESSIONAL EXPERIENCE:
Confidential
Data Analyst, Hicksville, NY
Responsibilities:
- Developed a custom ETL workflow using Hive and Python for creating flat files according to the requirements of the online ecommerce portals like Amazon, Groupon, Wayfair, Walmart, eBay, etc.
- Performed the various transformations on the data using Pig and Python.
- Wrote complex SQL queries to join the tables and retrieve data for inventory management and inventory health reports.
- Implemented the partitions and Buckets on Hive tables to improve the performance
- Lead the Amazon US and International FBA inventory removal team. Reduced inventory processing time by merging amazon removal shipping details with internal database.
- Lead the project for optimizing shipping times and cost. Utilized Spark’sGraphx library to determine geographical sales trends.
- Optimized our website’s recommender system by using machine learning and AWS tools.
- Created weekly and monthly dashboards on KPI monitoring, email and social media campaign reports, inventory health, trend report and sales reports using tableau.
- Developed web scarping programs using python beautifulsoup library tocollect trend data and competitor data for price point analysis and search engine optimization.
Tools: and Applications utilized: AWS, Lambdas, SageMaker, Hive, Spark, Tableau, ETL, Python, SQL, Pig
Confidential
Research Assistant, Hempstead, NY
Responsibilities:
- Gained experience in setting up a 6 node Hadoop cluster and performed various projects under the guidance of Dr. Kim.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Successfully completed projects on city data for fire incidents and comparison of weather data and road accidents.
- Gained hands on experience for implementing machine learning algorithms and various other statistical methods on spark and r studio for faster data processing.
- Twitter sentiment analysis using pyspark and streaming using textblob, tweepy and scikit-learn libraries.
- Developed a smart parking android application which allowed social interaction using MIT appinventor.
Tools: and Applications utilized: Ambari, Hortonworks Sandbox, Spark, Python, SQL, Sqoop
Confidential
Software Developer
Responsibilities:
- Involved in all phases of SDLC including Analysis and requirement determination, estimation, design, development, and Unit testing.
- Analyzing the Business requirements, understanding the specifications and scope of the requirements.
- Responsible for the project planning, providing estimations and the development of WBS.
- Designing and documenting technical solutions for the business requirements with the help of UML diagrams and state diagrams.
- Developing client customized interfaces for various clients using CSS and JavaScript.
- Performing the code review for peers and maintaining of the code repositories using GIT.
