Big Data Engineer Resume
3.00/5 (Submit Your Rating)
SUMMARY
- Designed and implemented data ingestion techniques for real time data coming from various data sources
- Data Collection and exploration (Python) + Data Visualization
- Experience in loading data coming from different data sources into HDFS and automate data ingestion and transformation jobs
- Performeddataextraction anddatawrangling using Pandas and Numpy modules in Python
- Programmed in Hive, Spark SQL and Python to streamline the incomingdataand build thedatapipelines to get the useful insights.
- Managed multiple tasks and worked under tight deadlines and in fast pace environment
- Possess good communication, interpersonal, analytical skills and a go - getter personality
TECHNICAL SKILLS
Analytical Tools: SQL, Jupyter Notebook, Tableau, Zeppelin, Graph Database,Talend, Tableau
Programming: Python - Data Manipulation, Numpy, Pandas, Matplotlib, Plotly
Big Data: Spark, Pig, Hive, Sqoop, HBase, Hadoop, HDFS, MapReduce
NoSQL: Cassandra, MangoDB
Methodologies: Agile and Waterfall model
Others: TWS, Shell Script
PROFESSIONAL EXPERIENCE
Confidential
Big Data Engineer
Responsibilities:
- Defining the metrics for the Big data analytics proof of concepts
- Defining the requirements for data lakes/pipe lines
- Efficiently handled periodic exporting of SQL data into Elasticsearch
- Creating the tables in Hive and integrating data between Hive &Spark
- Developed python scripts to collect data from source systems and store it on HDFS to run analytics
- Created Hive Partitioned and Bucketed tables to improve performance
- Created Hive tables with User defined functions
- Involved in code review and bug fixing for improving the performance
- Worked on the core and Spark SQL modules of Spark extensively using programming languages likeScala, R, Python
- Perform extensive studies of different technologies and capture metrics by running different algorithms
- Converting the SAS algorithms into different technologies
- Design and implement data ingestion techniques for real time data coming from various source systems
- Defining the data layouts and rules and after consultation with ETL teams
- Worked in aggressive AGILE environment and participated in daily Stand-ups/Scrum Meetings
Confidential
Big Data Engineer
Responsibilities:
- Data analysis using open source tools
- Design and implement data ingestion techniques for real time data coming from various source systems
- Creation of regulatory reports and analysis. Defining the data streams
- Defining the data layouts and rules and after consultation with ETL teams
- Importing and exporting data into HDFS and Hive using Sqoop.
- Written Hive queries for data analysis to meet the Business requirements.
- Experience in managing and reviewing Hadoop log files.
- Worked in aggressive AGILE environment and participated in daily Stand-ups/Scrum Meetings
Confidential
Python Developer
Responsibilities:
- Involved in the Design, development, test, deploy and maintenance of the website
- Data Analysis using python libraries
- Design and develop the Python code as per user requirements
- Debugging Software for Bugs. Environment: Python, DOM, HTML, CSS, SQL, PLSQL, Oracle and Windows
Environment: Python, DOM, HTML, CSS, SQL, PLSQL, Oracle and Windows
Confidential
Python Developer
Responsibilities:
- Involved in various phases of Software Development Life Cycle (SDLC) such as requirements gathering
- Modelling, analysis, design and development.
- Design, develop, test, deploy and maintain the website.
- Designed and developed data management system using MySQL.
- Rewrite existing Python/Django module to deliver certain format of data.
- Wrote python scripts to parse XML documents and load the data in database.
- Generated Use case diagrams, Activity flow diagrams, Class diagrams and Object diagrams in the design
Environment: Python, Shell scripting, PL/SQL, Oracle, SVN, Quality Center, Windows, Perl.
