Senior Hadoop Developer Resume
Columbus, OhiO
SUMMARY:
- Around 8 years of IT industry experience encompassing wide range of skill set in Big Data technologies and Java/J2EE technologies.
- 3+ years of experience in working with Big Data Technologies on systems which comprises of massive amount of data running in highly distributive mode in Cloudera, Horton works Hadoop distributions.
- Strong knowledge on Hadoop eco systems including HDFS, Hive, Oozie, HBase, Pig, Sqoop, Zookeeper, Flume, Kafka, MR2, Yarn, Spark etc.
- Excellent knowledge on Hadoop architecture 1.0 and 2.0 as in HDFS, Job Tracker, Task Tracker, Name Node, Data Node, Resource Manager, Node Manager and MapReduce programming paradigm.
- Good understanding of Data Replication, HDFS Federation, High Availability, Rack Awareness Concepts.
- Extensive Knowledge on developing Spark Streaming jobs by developing RDD’s (Resilient Distributed Datasets) and used pyspark and spark - shell accordingly.
- Experience on developing JAVA MapReduce jobs for data cleaning and data manipulation as required for the business.
- Executed complex HiveQL queries for required data extraction from Hive tables and written Hive UDF’s as required.
- A strong ability to prepare and present data in a visually appealing and easy to understand manner using Tableau, Excel etc.
- Having a great knowledge on Denodo. It performs many of the same transformation and quality functions as traditional data integration (Extract-Transform-Load (ETL), data replication, data federation, Enterprise Service Bus (ESB), etc.)
- Experience working with Cassandra and NoSQL database including MongoDB and Hbase.
- Managing and scheduling batch Jobs on a Hadoop Cluster using Oozie.
- Experience in managing and reviewing Hadoop Log files.
- Used Zookeeper to provide coordination services to the cluster.
- Hands on dealing with log files to extract data and to copy into HDFS using flume.
- Experience in analyzing data using Hive, Pig Latin, and custom MR programs in Java.
- Good understanding of different file formats like JSON, Parquet, Avro, ORC, Sequence, XML etc.
- Extensive experience on JAVA/J2EE technologies like Hibernate, Spring MVC
- Expertise in Core Java, data structures, algorithms, Object Oriented Design (OOD) and Java concepts such as OOPS Concepts, Collections Framework, Exception Handling, I/O System and Multi-Threading.
- Extensive experience on working with Soap and Restful web services.
- Extensive experience in working with Oracle, MS SQL Server, DB2, MySQL RDBMS databases.
- Experienced in working in SDLC, Agile and Waterfall Methodologies.
- Hands on experience with message brokers such as Apache kafka and RabbitMQ.
- Build and configured Apache TEZ on Hive and PIG to achieve better responsive time while running MR and Spark Jobs.
- Experience working with structured, semi-structured and unstructured data sets including social, web logs and real time data feeds.
- Experience in working with Health Care, and ecommerce industries.
- Ability to meet deadlines without compromising in delivering right output.
- Excellent Communication skills, Interpersonal skills, problem solving skills and a team player.
- Ability to quickly adapt new environment and technologies.
TECHNICAL SKILLS:
Big Data Frameworks: Hadoop, Spark, Scala, Hive, Kafka, AWS, Cassandra, HBase, Flume, Pig, Sqoop, MapReduce, Cloudera, Mongo DB.
Languages: Core Java, Scala, Python, SQL, Shell Scripting
Big data Distribution: Cloudera, Amazon EMR
Cloud: Amazon web services, Talend
Front End Technologies: JSP, HTML5, Ajax, JQuery and XML
Servers: CherryPy, Web Sphere and Tomcat.
Web and Enterprise Technologies: Servlet and Spring
Persistence: Hibernate
Visualization Tool: Apache Zeppelin & Matplotlib
Databases: Oracle, SQL Server
LINUX, Windows, MS: DOS
Control: M, Autosys
Query Tools: TOAD, SQL Navigator.
PROFESSIONAL EXPERIENCE:
Confidential, Columbus, Ohio
Senior Hadoop Developer
Responsibilities:
- Design data ingestion and integration process using SQOOP, Shell Scripts & Pig, with Hive.
- Adding and Decommissioning Hadoop Cluster Nodes Including Balancing HDFS block data.
- Implemented Fair schedulers on the Resource Manager to share the resources of the Cluster for the MRv2 jobs given by the users.
- Worked with the systems engineering team to propose and deploy new hardware and software environments required for Hadoop and to expand existing environments.
- Perform investigation and migration from MRv1 to MRv2.
- Developed Pyspark code to read data from Hive, group the fields and generate XML files.Enhanced the Pyspark code to write the generated XML files to a directory to zip them to CDAs
- Worked with Big Data Analysts, Designers and Scientists in troubleshooting MRv1/MRv2 job failures and issues with Hive, Pig, Flume, and Apache Spark.
- Utilized Apache Spark for Interactive Data Mining and Data Processing.
- Accommodate load in its place before the data is analyzed using Apache Kafka with its fast, scalable, fault-tolerant system.
- A analyzed the SQL scripts and designed the solution to implement using Pyspark.
- Configuring Sqoop to import and export data from HDFS to RDBMS and vice-versa.
- Handle the data exchange between HDFS & Web Applications and databases using Flume and Sqoop.
- Used Hive and created Hive tables involved in data loading.
- Extensively involved in querying using Hive, Pig.
- Developed open source Impala/Hive Liqui base plug-in to schema migration in CI/CD pipelines.
- Involved in writing custom UDF's for extending Pig core functionality.
- Involved in writing custom MR jobs which utilize Java API.
- Familiarity with NoSQL databases including Hbase and Cassandra.
- Implemented Cassandra connection with the Resilient Distributed Datasets.
- Design and develop JAVA API (Commerce API) which provides functionality to connect to the Cassandra through Java services.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Setup automated processes to analyze the System and Hadoop log files for predefined errors and send alerts to appropriate groups.
- Setup automated processes to archive/clean the unwanted data on the cluster, on Name node and Standby node.
- Created Gradle builds to build and deploy Spring Boot microservices to internal enterprise Docker registry.
- Created Maven builds to build and deploy Spring Boot microservices to internal enterprise Docker registry.
- Involved in Analyzing system failures, identifying root causes, and recommended course of actions. Documented the systems processes and procedures for future references.
- Supported technical team members in management and review of Hadoop log files and data backups.
- Designed target tables as per the requirement from the reporting team and designed Extraction, Transformation and Loading (ETL) using Talend.
- Implemented File Transfer Protocol operations using Talend Studio to transfer files in between network folders.
- Participated in development and execution of system and disaster recovery processes.
- Experience with cloud AWS and service like EC2, ELB, RDS, Elasti Cache, Route53, EMR.
- Hands on experience in cloud configuration for Amazon web services (AWS).
- Hands on experience with container technologies such as Docker, embed containers in existing CI/CD pipelines.
- Set up independent testing lifecycle for CI/CD scripts with Vagrant and Virtual box.
Environment: Hadoop, MapReduce2, Hive, Pig, HDFS, Sqoop, Oozie, Microservices, Talend, Pyspark, CDH, Flume, Kafka, Spark, HBase, Zookeeper, Impala, LDAP, NoSQL, MySQL, Info bright, Linux, AWS, Ansible, Puppet, AWS, Chef.
Confidential, Sacramento, CA
Hadoop Developer
Responsibilities:
- Handled importing of data from various data sources, performed data control checks using Spark and load data into HDFS.
- Developed several Restful web services supporting JSON to perform tasks
- Built real time pipeline for streaming data using Kafka and Spark Streaming.
- Spark Streaming collects this data from Kafka in near-real-time and performs necessary transformations and aggregation on the fly to build the common learner data model and persists the data in HBase.
- Developed Spark scripts by using Scala shell commands as per the requirement.
- Load the data into Spark RDD and performed in-memory data computation to generate the output response.
- Enhanced the Pyspark code to replace spark with Impyla. Performed installation for Impyla on the Edge node
- Experienced in implementing Spark RDD transformations, actions to implement business analysis.
- Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with HDFS reference tables and historical metrics.
- Experienced on loading and transforming of large sets of data from Cassandra source through Kafka and placed in HDFS for further processing.
- Implemented Cassandra connector for Spark 1.6.1
- Written customized Hive UDFs in Java where the functionality is too complex.
- Designed HBase schema to avoid Hot spotting and exposed data from tables to REST API on UI.
- Used Flume to collect, aggregate, and store the log data from different web servers.
- Created business data reports using Spark SQL.
- Implemented Spark using Scala and Spark SQL for faster testing and processing of data.
- Processed the streaming data using Kafka, integrating with Spark Streaming API.
- Worked on the Spark SQL for analyzing the data .
- Used Sqoop to import the data from RDBMS to Hadoop Distributed File System (HDFS) and later analyzed the imported data using Hadoop Components.
- Designed and Developed Hive managed/external tables using Struct, Maps and Arrays using various storage formats.
- Implemented various performance techniques ( Partitioning , Bucketing ) in hive to get better performance.
- Imported Hive tables into Impala for generating reports using Tableau .
- Developed workflows using Oozie to automate the tasks of loading the data into HDFS
- Used Oozie for automating the end to end data pipelines and Oozie coordinators for scheduling the work flows.
Environment : Hadoop, Sqoop, Hive, HDFS, YARN, Pyspark, Zookeeper, HBase, Apache Spark, Scala, Kafka, Oracle, Java, Spring IOC, Restful web service.
Confidential, Atlanta, GA
Hadoop/Big Data Engineer
Responsibilities:
- Configured Kafka/Flume ingestion pipeline to transmit the logs from web server to the Hadoop.
- Used interceptors with RegEx as part of flume configuration to eliminate the chunk from logs and dump the rest into HDFS.
- Used the Avro SerDe's for serialization & de-serialization of log files at different flume agents
- Created Pig Latin scripts for the duplication of the log files if any due to flume agent crash.
- Involved in partitioning the raw data, processed data each by day using one level partitioning schemes.
- Created the external tables in Hive based on the processed data obtained from Spark.
- Ingested the secondary data from systems like CRM, CPS, ODS using Sqoop and correlated this data with log files providing the platform for data analysis.
- Performed basic aggregations like count, average, sum, distinct, max, min on the existing hive tables using impala to determine Average Hit rates, Miss rates, Bounce rates etc.
- Persisted the processed data in columnar databases like HBASE and provided the platform for analytics using BI tools, analytical tools like R, machine learning such as Mahout.
- Involved in running and orchestrating the entire flow daily using Oozie jobs.
- Able to tackle the problems and accomplished the tasks which should be done during the sprint.
Environment: - Flume 1.5.2, Sqoop1.4.6, HDFS2.6.0, Hadoop2.6.0, Hive0.14.0, Hbase0.98.0, Impala2.1.0, Pig 0.14.0, Oozie 4.1.0
Confidential - Cherry Hill, NJ
Hadoop Developer
Responsibilities:
- Implemented Web service calls for Different integrations.
- Responsible for building scalable distributed data solutions using Hadoop.
- This project will download the data that was generated by sensors from the cars activities, the data will be collected in to the HDFS system online aggregators by Kafka.
- Experience in creating Kafka producer and Kafka consumer for Spark streaming which gets the data from different learning systems of the patients.
- Spark Streaming collects this data from Kafka in near-real-time and performs necessary transformations and aggregation on the fly to build the common learner data model.
- Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing.
- Experience in AWS to spin up the EMR cluster to process the huge data which is stored in S3 and push it to HDFS. Implemented automation and related integration technologies with Puppet.
- Implemented Spark SQL to access hive tables into spark for faster processing of data.
- Involved in Converting Hive/SQL queries into Spark transformations using Spark RDD, Scala.
- Interacting with Cloudera support and log the issues in Cloudera portal and fixing them as per the recommendations.
- Upgraded the Cloudera Hadoop ecosystems in the cluster using Cloudera distribution packages.
- Debug and solve the major issues with Cloudera manager by interacting with the Cloudera team.
- Worked on the proof-of-concept for Apache Hadoop 1.20.2 framework initiation.
- Used Apache Oozie for scheduling and managing the Hadoop Jobs. Extensive experience with Amazon Web Services (AWS)
- Developed Python/Django application for Google Analytics aggregation and reporting.
- Developed and updated social media analytics dashboards on regular basis.
- Good understanding of NoSQL databases such as HBase, Cassandra and MongoDB.
- Supported Map Reduce Programs running on the cluster and wrote custom Map Reduce Scripts for Data Processing in Java.
- Monitored workload, job performance and capacity planning using Cloudera Manager.
- Worked on migrating PIG scripts and Map Reduce programs to Spark Data frames API and Spark SQL to improve performance Involved in moving all log files generated from various sources to HDFS for further processing through Flume and process the files by using some piggy bank.
- Used Flume to collect, aggregate and store the web log data from different sources like web servers, mobile and network devices and pushed into HDFS. Used Flume to stream through the log data from various sources.
- Using Avro file format compressed with Snappy in intermediate tables for faster processing of data. Used parquet file format for published tables and created views on the tables.
- Created sentry policy files to provide access to the required databases and tables to view from impala to the business users in the dev, test and prod environment.
- Used Amazon Kinesis Data Streams to build custom applications that analyze data streams using popular stream processing frameworks.
- Used Amazon Kinesis Data Analytics to analyze data streams with SQL
- Good understanding of ETL tools and how they can be applied in a Big Data environment.
Environment:: Hadoop, Map Reduce, Cloudera, Spark, Kafka, HDFS, Hive, Pig, Oozie, Scala, Eclipse, Flume, Kinesis, Oracle, UNIX Shell Scripting
Confidential
Java Developer
Responsibilities:
- Involved in the development of use case documentation, requirement analysis, and project documentation.
- Developed and maintained Web applications as defined by the Project Lead.
- Developed GUI using JSP, JavaScript, and CSS.
- Used MS Visio for creating business process diagrams.
- Developed Action Servlet, Action Form, Java Bean classes for implementing business logic for the struts Framework.
- Developed Servlets and JSP based on MVC pattern using struts Action framework.
- Developed all the tiers of the J2EE application. Developed data objects to communicate with the database using JDBC in the database tier, implemented business logic using EJBs in the middle tier, developed Java Beans and helper classes to communicate with the presentation tier which consists of JSPs and Servlets.
- Used AJAX for Client side validations.
- Applied annotations for dependency injection and transforming POJO/POJI to EJBs.
- Developed persistence layer modules using EJB Java Persistence API (JPA) annotations and Entity manager.
- Involved in creating EJBs that handle business logic and persistence of data.
- Developed Action and Form Bean classes to retrieve data and process server side validations.
- Designed various tables required for the project in Oracle database and used Stored Procedures in the application. Used PL SQL to create, update and manipulate tables.
- Used IntelliJ as IDE and Tortoise SVN for version control.
- Involved in impact analysis of Change requests and Bug fixes.
Environment:: Java 5, Struts, PL/SQL, Oracle, EJB, IntelliJ, Tortoise SVN, MS Visio, Firebug, Apache Tomcat, JSP, Java Script, CSS.
