We provide IT Staff Augmentation Services!

Big Data Engineer Resume

2.00/5 (Submit Your Rating)

Long Beach, CA

PROFESSIONAL SUMMARY

  • 8+ years of experience in Architect, Analysis, Design, Development, Testing, Implementation, Maintenance and Enhancements on various IT Projects and experience in Big Data in implementing end - to-end Hadoop solutions.
  • Excellent experience in installing, configuring and using Apache Hadoop ecosystem components like Hadoop Distributed File System (HDFS), MapReduce, PIG, HIVE, HBASE, Apache Crunch, ZOOKEEPER, SQOOP, Hue, Spark, Storm, Kafka Solr, NiFi, Git, Maven, AVRO, JSON and CHEF.
  • Expertise in depth understanding/knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, MRv1 and MRv2 (YARN).
  • Experienced with data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling and data mining, machine learning and advanced data processing.
  • Expertise in writing Hadoop Jobs to analyze data using MapReduce, Apache Crunch, Hive, Pig and Solr in java.
  • Experienced working with JIRA for project management, GIT for source code management, JENKINS for continuous integration and Crucible for code reviews.
  • Expertise in writing Apache Spark streaming API on Big Data distribution in the active cluster environment.
  • Experienced on implementation of a log producer in Scala that watches for application logs, transform incremental log and sends them to a Kafka and Zookeeper based log collection platform.
  • Experienced in working with Flume to load the log data from multiple sources directly into HDFS.
  • Excellent knowledge in building and scheduling Big Data workflows with the help of OOZIE and Auto-sys.
  • Experienced in importing and exporting data from the different Data sources like (Teradata and DB2) using Sqoop from HDFS to Relational Database Systems (RDBMS) and vice-versa and load into partitioned Hive tables.
  • Designed and implemented Hive and Pig UDF's using Python for evaluation, filtering, loading and storing of data.
  • Experience working with Neo4j which is a NoSQL graph database.
  • Developed Simple to complex Map/reduce streaming jobs and wrote REST API’s using Python languages that are implemented using Hive and Pig.
  • Expertise in NOSQL database, Hbase, MongoDB and Cassandra.
  • BigData Proficient in using Cloudera Manager, an end-to-end tool to manage Hadoop operations.
  • Experienced working with Horton works Distribution and Cloudera Distribution.
  • Expertise in Amazon AWS concepts like EMR and EC2 which provides fast and efficient processing.
  • Experience in Elastic Search (used for Faster Indexing), Kibana (Creating Dashboards), Splunk (Log Analysis and Dashboards).
  • Expertise in core Java, J2EE, Multithreading, JDBC, Hibernate, spring, Confidential Scripting and proficient in using Java API’s for application development.
  • Usage of Gradle to build and automate Hadoop jobs for containers.
  • Usage of Docker which combines an easy-to-use interface to Linux containers with easy-to-construct image files for those containers.
  • Experience with SQL RDBMS like SQL Server, Oracle and My SQL or MPP databases like Vertica, Teradata and Netezza.
  • Expertise in core Java, J2EE, Multithreading, Spring, Confidential Scripting.
  • Excellent working experience in Scrum / Agile framework and Waterfall project execution methodologies.

TECHNICAL SKILLS

Hadoop Technologies: Apache Hadoop, CDH 4, CDH 5 & HDP 2.4.2

Hadoop Ecosystem: Hive, Pig, Sqoop, Flume, Zookeeper, Oozie.

Streaming Technologies: Spark, Kafka, Storm.

AWSS3: EC2, EMR, IAM, Redshift, Snowflake.

Java/J2EE Technologies: Core Java, Data Structures, Multithreading.

NOSQL Databases: Hbase.

Programming Languages: Java, Linux Confidential scripting, Scala, Python.

Web Technologies: HTML, CSS, JavaScript, AJAX, JSP, DOM, XML

DatabasesMy: SQL, SQL, Oracle, SQL Server, DB2, PL/SQL.

Application Servers: Web Logic, Web Sphere, JBoss.

Software Engineering: UML, Object Oriented Methodologies, Scrum, Agile methodologies ETLTalend.

Operating Systems: Windows 95/98/2000/XP, MAC OS, UNIX, LINUX.

IDE Tools: Eclipse, IntelliJ IDEA

PROFESSIONAL EXPERIENCE

Confidential, Long Beach, CA

Big Data Engineer

Responsibilities:

  • Loading the data from the different Data sources like (Teradata and DB2) into HDFS using SQOOP and load into Hive tables, which are partitioned.
  • Developed Hive UDF's to bring all the customers email id into a structured format.
  • Developed bash scripts to bring the Tlog files from ftp server and then processing it to load into hive tables.
  • Strom and Kafka queue to send push messages to mobile devices
  • Inserted Overwriting the HIVE data with HBase data daily to get fresh data every day and used Sqoop to load data from DB2 into HBASE environment.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Scala and have a good experience in using Spark- Confidential and Spark Streaming.
  • Designed, developed and maintained Big Data streaming and batch applications using Storm.
  • Experienced in Core Java with strong understanding of Multithreading, Collections, Concurrency, and Exception handling concepts, Object-oriented analysis, design, and development.
  • Import millions of structured data from relational databases using Sqoop import to process using Spark and stored the data into HDFS in CSV format.
  • Created Hive, Phoenix, HBase tables and HBase integrated Hive tables as per the design using ORC file format and Snappy compression.
  • Developed UDF's using both DataFrames/ SQL and RDD in Spark for data Aggregation queries and reverting into OLTP through Sqoop.
  • All the bash scripts are scheduled using Resource Manager Scheduler.
  • Developed Oozie Workflows for daily incremental loads, which gets data from Teradata and then imported into hive tables.
  • Experience on implementation of a log producer in Scala that watches for application logs, transform incremental log and sends them to a Kafka and Zookeeper based log collection platform.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDD, Scala and Python.
  • Worked on Sequence files, RC files, Map side joins, bucketing, partitioning for Hive performance enhancement and storage improvement.
  • Sqoop jobs, PIG and Hive scripts were created for data ingestion from relational databases to compare with historical data.
  • Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
  • Developed pig scripts to transform the data into structured format and it are automated through Oozie coordinators.
  • Used Splunk to captures, indexes and correlates real-time data in a searchable repository from which it can generate reports and alerts.
  • Fix the code review comments; Build the Jenkins and support for the code deployment into the production. Fix the postproduction defects to perform the Map/Reduce code to work as expected.

Environment: Hadoop, HDFS, Spark, Strom, Kafka, Map Reduce, Hive, Pig, Sqoop, Oozie, DB2, Java, Python, Splunk, UNIX Confidential Scripting.

Confidential, Detroit, MI

Lead Hadoop Developer

Responsibilities:

  • Created the Spark Streaming code to take the source files as input.
  • Developed simple to complex Map Reduce job using Hive.
  • Used Sqoop extensively to import data from RDMS sources into HDFS. Performed transformations, cleaning and filtering on imported data using Hive, Map Reduce, and loaded final data into HDFS
  • Provisioning of Cloudera Director AWS instance and adding Cloudera manager repository to scale up Hadoop Cluster in AWS.
  • Encryption Mechanisms using Python.
  • Involved in loading data from UNIX file system to HDFS using Flume and Kettle and HDFS API.
  • Involved in managing and reviewing Hadoop log files.
  • Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with reference tables and historical metrics.
  • Involved in running Hadoop streaming jobs to process terabytes of text data.
  • Developed HIVE queries for the analysts.
  • Experience working with NoSQL databases such as HBase and Cassandra.
  • Used spark machine learning technique implemented in Scala.
  • Developing the ETL mappings for XML, .CSV, .TXT sources and loading the data from these sources into relational tables with Talend ETL.
  • Worked on using Talend Integration suite and created many jobs on Talend
  • Developed the technical strategy for Spark integrated for pure streaming and more general data-computation needs.
  • Configured Spark Streaming to receive real time data from the Kafka and store the stream data to HDFS.
  • Developed complex Talend jobs mappings to load the data from various sources using different components.
  • Created validation and error calculation mapplets using Talend ETL.
  • Developed the technical strategy for Spark integrated for pure streaming and more general data-computation needs.
  • Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
  • Exported the result set from HIVE to MySQL using Kettle (Pentaho data-integration tool).
  • Used Zookeeper for various types of centralized configurations.
  • Worked on implementation of a log producer in Scala that watches for application logs, transform incremental log and sends them to a Kafka and Zookeeper based log collection platform.
  • Designed and Developed ETL jobs using Talend Big Data ETL.
  • Experience Managing and Leading Offshore team.
  • Lead the team and conducted knowledge sessions with the team of 8.
  • Implemented Fair schedulers on the Job tracker to share the resources of the Cluster for the Map
  • Implemented Spark using Scala and SparkSQL for faster testing and processing of data.
  • Monitor System health and logs and respond accordingly to any warning or failure conditions.

Environment: Hadoop (Cloudera), HDFS, Map Reduce, Kafka, Hive, Scala Pig, Sqoop, Oozie, AWS, Solaris, DB2, SparkSQL, Spark Streaming, Spark, Python, Django, UNIX Confidential Scripting.

Confidential, Houston, TX

Big Data Engineer

Responsibilities:

  • Extracted files from DB2 through Kettle and placed in HDFS and processed.
  • Analyzed large data sets by running Hive queries and Pig scripts.
  • Developed the Sqoop scripts to make the interaction between Hive and vertica Database.
  • Involved in creating Hive tables, and loading and analyzing data using hive queries.
  • Developed Simple to complex MapReduce Jobs using Hive and Pig.
  • Involved in running Hadoop jobs for processing millions of records of text data.
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
  • Developed multiple MapReduce jobs in java for data cleaning and pre-processing.
  • Involved in unit testing using MR unit for MapReduce jobs.
  • Involved in loading data from LINUX file system to HDFS.
  • Loading data from multiple sources on AWS S3 cloud storage.
  • Experienced in running Hadoop streaming jobs to process terabytes of xml format data.
  • Load and transform large sets of structured, semi structured data.
  • Assisted in exporting analyzed data to relational databases using Sqoop.
  • Created and maintained Technical documentation for launching HADOOP Clusters and for executing Hive queries and Pig Scripts.

Environment: Hadoop, HDFS, Pig, Hive, MapReduce, AWS S3, Sqoop, LINUX, MRUnit and Big Data.

Confidential, Houston, TX

Hadoop Developer

Responsibilities:

  • Setup and benchmarked Hadoop /HBase clusters for internal use.
  • Developed Java MapReduce programs for the analysis of sample log file stored in cluster.
  • Developed MapReduce programs to cleanse the data in HDFS obtained from heterogeneous data sources to make it suitable for ingestion into Hive schema for analysis
  • Developed multiple scripts for analyzing data using Hive and Pig and integrating with HBase.
  • Used Sqoop to import data into HDFS and Hive from other data systems.
  • Created reports for the BI team using Sqoop to export data into HDFS and Hive.
  • Migration of ETL processes from Oracle to Hive to test the easy data manipulation.
  • Data was pre-processed and fact tables were created using HIVE.
  • The resulting data set was exported to SQL server for further analysis.
  • Create Hive scripts to extract, transform, load (ETL) and store the data.
  • Automated all the jobs from pulling data from databases to loading data into SQL server using Confidential scripts.

Environment: Apache Hadoop, HDFS, Java, MapReduce, Hive, PIG, Sqoop, SQL.

Confidential

Java Developer

Responsibilities:

  • Involved in SDLC Requirements gathering, Analysis, Design, Development and Testing of application developed using AGILE methodology.
  • Actively participated in Object Oriented Analysis Design sessions of the Project, which is based on MVC Architecture using Spring Framework.
  • Involved in Daily Scrum meetings, Sprint planning and estimation of the tasks for the user stories, participated in retrospective and presenting Demo at end of the sprint.
  • Experience in writing PL/SQL Stored procedures, Functions, Triggers, Oracle reports and Complex SQL’s.
  • COTS Evaluation and implementation for reporting tool, that resulted in choosing Business Objects.
  • Experience in developing Unit testing & Integration testing with unit testing frameworks like JUnit, Mockito, TestNG, Jersey Test and Power Mocks.
  • Designed and developed entire application implementing MVC Architecture.
  • Developed frontend of application using Bootstrap (Model, View, and Controller), Java Script, and Angular.js framework.
  • Used Spring framework for implementing IOC/JDBC/ORM, AOP and Spring Security.
  • Involved in Java, J2ee, Spring 4.0, Restful Web Services, WebSphere 5.0/6.0 in a fast-paced development environment.
  • Proficient in developing applications having exposure to Java, JSP, UML, Servlets, Struts, Swing DB2, Oracle (SQL, PL/SQL), HTML, Junit, JSF, Java Script, CSS.
  • Proactively found the issues and resolved them.
  • Established efficient communication between teams to resolving the issues.
  • Gave an innovative for logging for all interdepends application.
  • Successfully delivered all product deliverables that resulted with zero defects.
Environment: JDK 1.7, Oracle 11g, Junit, MySQL, JSP, UML, JUNIT, Angular.js, jQuery, COTS, Hibernate 4.0, Spring, Struts, WebSphere 5.0/6.0, SOAP, JSF, HTML, CSS, Web Services (SOAP, RESTFUL 4.0)

We'd love your feedback!