Hadoop Developer Resume
New York, NY
SUMMARY:
-
6+ years of professional experience in IT industry as a software Engineer with a background in analysis, development, integration and testing of applications.
- Worked in various domains including Media, Finance, Insurance and E - commerce
- 3+ years worked as a dedicated professional Big Data Engineer with a solid background in Hadoop ecosystem: HDFS, MapReduce, Spark, Hive, Kafka, HBase, Pig, Sqoop, Flume and Zookeeper
- Have a deep understanding of workload management, schedulers, scalability and distributed platform architectures
- Experience in Spark programing with Scala for high-volume data processing
- Experience in collecting, transferring, aggregating and processing large amounts of streaming data using Kafka, Flume, Spark Streaming
- Experience in writing Pig Latin scripts and HiveQL Queries for preprocessing and analyzing large volumes of data
- Experience in writing MapReduce programs with Java for data processing in Hadoop
- Experience in importing and exporting buck of data using Sqoop from HDFS/Hive/HBase to RDBMS
- Experience in RDBMS including Oracle, MySQL
- Experience in developing scalable solutions using NoSQL databases including Cassandra, HBase and MongoD B
- Knowledge of data serialization and familiar with data format s including XML, JSON, sequenceFile, Avro and Parquet
- Experience on commercial distribution of Hadoop including Horton Works HDP and Cloudera CD H
- Experience in Hadoop cluster administration & performance tuning
- Experience in all the phases of Data warehouse life cycle involving requirement analysis, design, coding, testing, and deployment
- Strong in Core Java, D ata S tructure and Algorithms, and Object-Oriented Design
- Experienced with front-end technology including HTML, CSS, Bootstrap, JavaScript, Ajax, JQuery and AngularJS
- Experience in using varieties web frameworks including Hibernate, Spring and Node.js
- Experience in Unit Testing with JUnit, Scala Test, Pytest
- Family with software development tools like Git, JIRA, Jenkins, SVN
- Expose to various software development methodologies like WaterFall, Spiral development and Agile/Scrum
- A good team-player, can work independently in a fast-paced multitasking environment, and a self-motivated learner
TECHNICAL SKILLS:
Hadoop/Spark Ecosystem \ Database: Hadoop 2.x, MapReduce, Spark 2.x, Pig 0.12, \ Oracle 11g, MySQL 5.x, HBase 0.98, \ Hive 0.14, Sqoop 1.4.6, Flume 1.6.0, Kafka 2.10, \ Cassandra 2.0, MongoDB 3.2\ Yarn, Mesos, Zookeeper.
Programming Language\ Web Development Framework: Java, Scala, Python, SQL, Unix/Bash shell, \ JQuery, Ajax, AngularJS, Bootstrap, Hibernate, \ C/C++, JavaScript, HTML, CSS, XML\ Spring, Node.js.
Operating System\ Cloud Platform: Linux, Mac OS, Windows\ Amazon Web Servies EC2/S3/EMR.
Environment: & Tools\ IDE: Git/Github, Agile/Scrum, SVN, JIRA, Jenkins\ IntelliJ IDEA, Eclipse.
P ROFESSIONAL EXPERIENCE
Confidential, New York, NY
Data Engineer
Responsibilities:
-
Developed Kafka producers and consumers efficient ingested data from different data sources and decoupling the data
- Developed Spark Streaming programs to process real time data from Kafk a, and process data with both stateless and stateful transformations
- Developed Spark programs with Scala and Spark SQL, and applied principles of functional programming to do batch process ing
- Built a Cassandra data model based on the data from Spark Streaming, and utilized CQL to partitioning and fast retrieving the data
- Stored both the raw data and computation results in the Cassandra for future decision support and BI analytics
- Configure d M esos to coordinate and support Kafka, Spark, Cassandra and HDFS
- Integra ted Flume with Kafka, and Worked on monitoring and troubleshooting the Kafka-Flume-HDFS data pipeline for real-time data ingestion in HDFS
- Performed unit testing using ScalaTest
- Used Git for version control and JIRA for project tracking
- Actively Participated in software development lifecycle including scope, design, implement, testing and code reviews.
- Involved in story-driven Agile development methodology and actively participated in daily Scrum meetings
Environment: Hadoop, Cloudera CDH, HDFS, Kafka, Spark, Spark Streaming, Spark S QL, Cassandra, CQL, Flume, Mesos, Zookeeper, ScalaTest, Git, JIRA
Confidential, New York, NY
Hadoop Developer
Responsibilities:
-
Applied Spark using Scala to do the data batch processing, and store the output in HBase for scalable storage and fast query
- D esigned and created of Hive tables and worked on various performance optimizations like Partition, Bucketing in Hive
- Implemented Hive custom UDF s and Analyzed large data sets by running HiveQL to achieve comprehensive data analysis
- M igrated of MapReduce jobs and Hive quer ies into Spark transformations and actions to improve the performance
- Utilized Sqoop to import and output data between Oracle database and HDFS
- Configure the Sqoop i Confidential emental import job for importing the updated input data
- Convert raw data with sequence data format, such as Avro, and Parquet to reduce data processing time and i Confidential ease data transferring efficiency through the network
- Involved in application performance tuning and troubleshooting
- Collaborate and tracking the work with Git and JIRA
- Actively participated and provided feedback constructively during daily Stand up meetings and weekly Iterative review meetings
Environment: Hadoop, HDFS, Sqoop, Hive, Spark, Scala, HBase, Git, JIRA, Agile
Confidential, New York, NY
Jr. Hadoop Developer
Responsibilities:
-
Experienced on loading and transforming of large sets of structured and semi structured data
- Created Hive tables, analyzed data with Hive Queries, and written Hive UDFs
- Experience in using Partitions, bucketing to create Hive tables for performance optimization
- Experience in writing Pig-Latin scripts for data preprocessing
- Migrated data between RDBMS and HDFS/Hive with Sqoop
- Experience in defining job flows and wrote simple to complex MapReduce jobs
- Cluster coordination services through Zookeeper
- Performed performance tuning and troubleshooting of MapReduce jobs by analyzing and reviewing Hadoop log files
- Involved in Unit testing using JUnit and MRUnit
- Involved in reviewing Functional requirements and designing solutions
- Documented systems processes and procedures for future references
- Involved in gathering the requirements, designing, development and testing
Environment: Hadoop, Map Reduce, HDFS, Java, Pig, Hive, Sqoop, Shell Scripting, Linux, Oracle 10g
Confidential, Hangzhou, CN
Java Developer
Responsibilities:
-
Involved in designing user screens and validations using HTML, jQuery, CSS3, JSP, Servlet
- Responsible for validation of Client interface JSP pages using jQuery form validations
- Used Spring Dependency Injection property to provide loose-coupling between layers
- Developed Web services (SOAP) for transmission of large blocks of XML data over HTTP
- Used Hibernate ORM framework with spring framework for data persistence and transaction management
- Involved in the implementation of Spring-Hibernate ORM and creating the Hibernate POJO Objects and mapped with Oracle database using Hibernate Annotations
- Wrote SQL queries, stored procedures, and triggers to perform back-end database operations
Environment: JavaEE, JSP, Spring, Hibernate, JDBC, CSS3, JavaScript, XML, Oracle 10g, Eclipse
Confidential, Hangzhou, CN
Jr. Java Developer
Responsibilities:
-
Developed front end payment interface with HTML, CSS, JavaScript, JSP, AJAX and jQuery
- Implemented data persistency using JDBC for database connectivity and Hibernate for database/java object mapping.
- Implemented Jersey JAX-RS to create RESTful Web Services
- Developed unit test cases with JUnit
- Wrote SQL for querying, inserting and managing the database
- Worked in Spring MVC Framework with Agile methodology
Environment: Spring MVC, MySQL, JSP, JDBC, CSS, JavaScript, AJAX, Hibernate 3, JUnit, RESTful, Agile
Confidential, Hangzhou, CN
Jr. Java Developer
Responsibilities:
-
Used Spring framework for dependency injection, transaction management
- Designed the user interfaces using JSP , HTML & CSS
- Implemented client side validation using JavaScript and Server side validation of form data using Struts validation framework
- Developed the application using Struts Framework to implement a MVC design approach
- Configured the Hibernate configuration files to persist the data to the Oracle Database
- Designed Integration test plan for testing of Integration of all use cases for a module
Environment: Java1.6/J2EE, JSP, HTML, CSS, JavaScript, Oracle 9i/10g, Hibernate, XML, Struts2, Spring
