We provide IT Staff Augmentation Services!

Sr. Hadoop Developer Resume

5.00/5 (Submit Your Rating)

Dover, NH

SUMMARY:

  • Over 8 years of experience in full Software Development Life Cycle (SDLC), AGILE Methodology and analysis, design, development, testing, implementation and maintenance in Hadoop, Data Warehousing, Linux and Java.
  • 4+ years of experience in providing solutions for Big data using Hadoop 2.x, HDFS, MR2, YARN, Kafka, PIG, Hive, Sqoop, HBase, Cloudera Manager, Zoo keeper, Oozie, Hue, CDH5 & HDP 2.x.
  • Experienced in Big data, Hadoop, NoSQL and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce2, YARN programming paradigm.
  • Good understanding of distributed systems, HDFS architecture, Internal working details of MapReduce and Spark processing frameworks. Worked on debugging, performance tuning of Hive Jobs.
  • Migrating tables from RC format to ORC and data induction and other customized file formats.
  • Extending Pig and Hive core functionality by writing customized User Defined Functions for analysis of data, file processing, by running PigLatinScripts.
  • Having experience in creating Hive internal/external Tables using shared Meta Store.
  • Written Sqoop Queries to import data into Hadoop from Teradata/SQLServer.
  • Extensive experience working with real time streaming applications and batch style large scale distributed computing applications, worked on integrating Kafka with NiFi and Spark.
  • Good conceptual understanding and experience in cloud computing applications using Amazon EC2, S3, EMR.
  • Proficient in applying performance tuning concepts to SQL Queries, Informatica Mappings, Session and workflow properties, and database.
  • Hands on experience with opens source monitoring tools Ambari and Cloudera Manager.
  • Worked on Producer API and created a custom partitioner to publish the data to the Kafka Topic. Worked on POC for streaming data using Kafka and spark streaming.
  • Implemented Kafka Customer with Spark - streaming and Spark SQL using Scala. Validated the Dstream and created generated new Dstream and saved the data in HDFS.
  • Involved in importing the real time data to Hadoop using Kafka and implemented the Oozie job for daily. Involved in developing Hive DDLs to create, alter and drop Hive tables and storm, & Kafka.
  • Used Apache Oozie for scheduling and managing the Hadoop Jobs
  • Developed Python/Django application for Google Analytics aggregation and reporting.
  • Having extensive knowledge of RDBMS such as Oracle, Microsoft SQL Server and MYSQL.
  • Good understanding of No SQL databases such as HBase, Cassandra and MongoDB.
  • Supported MapReduce Programs running on the cluster and wrote custom MapReduce Scripts for Data Processing in Java.
  • Experience with operating ETL processes and data pipelines to build large, complex data sets.
  • Encountered in developing Spark jobs using Scala in test environment for faster data processing and used Spark SQL for querying.
  • Knowledge with installation and configuring Hadoop based monitoring tools (NagiOS/ Ganglia) is a plus.
  • Capable ofoading from disparate data sets - Familiarity with data loading tools like Flume, Sqoop.
  • Experience in Web Services using XML, HTML and SOAP.
  • Diverse experience in utilizing Java tools in business, Web, and client-server environments including Java Platform, J2EE, EJB, JSP, Java Servlets, Junit, Java database Connectivity (JDBC) technologiesand application servers like Web Sphere and Weblogic.
  • Familiarity in working with popular frameworks likes Struts, Hibernate, SpringMVC and AJAX.

TECHNICAL SKILLS:

Big data/Hadoop Ecosystem: HDFS, MR2, HIVE, PIG, HBase, Sqoop, Flume, Oozie, Storm,Airflow and Avro

Java / J2EE Technologies: Core Java, Servlets, JSP, JDBC, XML, REST, SOAP, WSDL

Programming Languages: C, C++, Java, Scala, SQL, PL/SQL, Linux shell scripts.

NoSQL Databases: MongoDB, Cassandra, Hbase

Database: Oracle 11g/10g, DB2, MS-SQL Server, MySQL, Teradata.

Web Technologies: HTML, XML, JDBC, JSP, JavaScript, AJAX, SOAP

Frameworks: MVC, Struts 2/1, Hibernate 3, Spring 3/2.5/2.

Tools: Used: Eclipse, IntelliJ, GIT, Putty, Winscp

Operating System: Ubuntu (Linux), Win 95/98/2000/XP, Mac OS, RedHat

ETL Tools: Informatica, pentaho.

Testing: Hadoop Testing, Hive Testing, Quality Center (QC)

Monitoring and Reporting tools: Ganglia, Nagios, Custom Shell scripts.

PROFESSIONAL EXPERIENCE:

Confidential, Dover,NH

Sr. Hadoop Developer

Responsibilities:

  • Creating end to end Spark-Solr applications using Scala to perform various data cleansing, validation, transformation and summarization activities according to the requirement
  • Implemented Moving averages, Interpolations and Regression analysis on input data.
  • Tuning spark application to improve performance. Worked collaboratively to manage build outs of large data clusters and real time streaming with Spark.
  • Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data.
  • Implement UDF s to consume complex Cassandra UDT.
  • Used Spark API over Horton works HDP 2.2 and Hadoop YARN to perform analytics on data in Hive
  • Experience with Horton works HDD2.2& Cloudera Manager Administration also experience in Installing, Updating Hadoop and its related components in Single node cluster as well as Multi node cluster environment using Ambari, Cloudera Manager 2.X,3.X,4.X, Horton works
  • Exploring with the Spark for improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Pair RDD's, YARN.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs and Scala.
  • Responsible for gathering the business requirements for the Initial POCs to load the enterprise data warehouse data to Greenplum databases.
  • Fixed a bug in the tls certificate checking code that prevented certificates with long chains to be used in golang.
  • Worked on Implementation of a log producer in Scala that watches for application logs transform incremental log and sends them to a Kafka and Zookeeper based log collection platform
  • Design and improve internal search engine using Big data and SOLR/Fusion.
  • Data migration from various data sources to SOLR via stages according to the requirement
  • Extensively worked on Jenkins for continuous integration and for End to End automation for all build and deployments.
  • Work with cross functional consulting teams within the data science and analytics team to design, develop, and execute solutions to derive business insights and solve clients' operational and strategic problems.
  • Converted Ant application into Gradle. Onsite-Offshore synchronization. Teams at both the ends should be well connected to have a smooth flow in the project and solve the roadblocks
  • Monitoring the ticketing tool for any tickets indicating an issue/incident reported and resolving with the appropriate fix in the project.
  • Worked on HBase to perform real time analytics and experienced in CQL to extract data from Cassandra tables.
  • Worked on Apache Nifi to Uncompress and move json files from local to HDFS
  • Migrated an existing on-premises application to AWS and services like EC2 and S3 for small data sets.
  • Used Cloud watch logs to move app logs to S3. Create alarms based on exceptions raised by applications.
  • Involved in loading and transforming large sets of Structured, Semi-Structured and Unstructured data and analyzed them by running Hive queries.
  • Work with Architecture and Development teams to understand usage patterns and work load requirements of new projects to ensure the Hadoop platform can effectively meet performance requirements and service levels of application.

Environment: Hadoop, AWS, Java, HDFS, MapReduce, Spark, Pig, Hive, Impala, Sqoop, Flume, Kafka, HBase, Oozie, Java, SQL scripting, Linux shell scripting, Eclipse and Cloudera.

Confidential, San Luis Obispo,CA

Sr. Hadoop Developer

Responsibilities:
  • Responsible for developing efficient MapReduce on AWS cloud programs for more than 20 years' worth of claim data to detect and separate fraudulent claims.
  • Worked with the advanced analytics team to design fraud detection algorithms and then developed MapReduce programs to efficiently run the algorithm on the huge datasets.
  • Ran data formatting scripts in Python and created terabyte csv files to be consumed by Hadoop MapReduce jobs.
  • Created ETL Mappings in Informatica for loading data into Data Warehouse
  • Performed data analysis, feature selection, feature extraction using Apache Spark Machine Learning streaming libraries in Python.
  • Involved in administration, installing, upgrading and managing CDH3, Pig, Hive & HBase.
  • Played a key-role is setting up a 50 node Hadoop cluster utilizing Apache Spark by working closely with the Hadoop Administration team.
  • Created Hive tables to store data into HDFS, loading data and writing hive queries that will run internally in map-reduce way.
  • Hands on experience in Spark and Spark Streaming creating RDD's, applying operations -Transformation and Actions. Extracting real time data using Kafka and spark streaming by Creating DStreams and converting them into RDD, processing it and stored it into Cassandra.
  • Uploaded and processed terabytes of data from various structured and unstructured sources into HDFS (AWS cloud) using Sqoop and Flume.
  • Involved in Cluster coordination services through Zookeeper.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Played a key role in installation and configuration of the various Hadoop ecosystem tools such as Solr, Kafka, Pig, HBase and Cassandra.
  • Implemented various hive optimization techniques like Dynamic Partitions, Buckets, Map Joins, Parallel executions in Hive.
  • Monitor datalake connectivity, security, performance and File system management
  • Conduct day-to-day administration and maintenance work on the datalake environment
  • Created data pipeline for different events of ingestion, aggregation and load consumer response data in AWS S3 bucket into Hive external tables in HDFS location to serve as feed for tableau dashboards.
  • Scheduled and executed workflows in Oozie to run Hive and Spark jobs.
  • Worked with Airflow (replaced the work of oozie).
  • Built centralized logging to enable better debugging using Elastic Search, Logstash and Kibana.
  • Efficiently handled periodic exporting of SQL data into Elastic search.
  • Worked on GitHub and Jenkins continuous integration tool for deployment of project packages.
  • Parse Json files through Spark core to extract schema for the production data using SparkSQL and Scala.
  • Involved in file movements between HDFS and AWS S3 and extensively worked with S3 bucket in AWS
  • Experienced with AWS services to smoothly manage application in the cloud and creating or modifying the instances.

Environment: Hadoop, HDFS, Pig, Hive, MapReduce, Sqoop, Kafka, CDH3, Cassandra, Python, Oozie, Java collection, Scala, AWS cloud, SQL, NoSQL, Bitbucket, Jenkins, HBase, Flume, spark, Solr, Zookeeper, ETL, Centos, Eclipse.

Confidential, Washington,DC

Hadoop Developer

Responsibilities:
  • Worked on analyzing Hadoop cluster and different big data analytic tools including Pig, Hbase database and Sqoop.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Implemented nine nodes CDH3 Hadoop cluster on Red hat LINUX.
  • Involved in loading data from LINUX file system to HDFS.
  • Worked on installing cluster, commissioning & decommissioning of datanode, namenode recovery, capacity planning, and slots configuration.
  • Created HBase tables to store variable data formats of PII data coming from different portfolios.
  • Implemented a script to transmit sysprin information from Oracle to Hbase using Sqoop.
  • Implemented best income logic using Pig scripts and UDFs.
  • Implemented test scripts to support test driven development and continuous integration.
  • Worked on tuning the performance Pig queries.
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
  • Wrote Teradata SQL scripts for trouble shooting and data comparison
  • Responsible to manage data coming from different sources.
  • Involved in loading data from UNIX file system to HDFS.
  • Load and transform large sets of structured, semi structured and unstructured data
  • Cluster coordination services through Zookeeper.
  • Experience in managing and reviewing Hadoop log files.
  • Installed Oozie workflow engine to run multiple Hive and pig jobs.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.

Environment: Hadoop, Spark Core, Spark-SQL, Spark-Streaming, MapReduce, HDFS, Hive, Java, Scala, Hue, SQL, Teradata, Pig, Sqoop, Tez, HBase, Zookeeper, PL/SQL, MySQL, DB2, Teradata.

Confidential, Dublin,OH

Hadoop Developer

Responsibilities:
  • Worked on analyzing Hadoop cluster and different big data analytic tools including Pig, Hbase database and Sqoop.
  • Worked on analyzing Hadoop cluster and different big data analytic tools including Pig, Hbase database and Sqoop.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Implemented nine nodes CDH3 Hadoop cluster on Red hat LINUX.
  • Involved in loading data from LINUX file system to HDFS.
  • Worked on installing cluster, commissioning & decommissioning of datanode, namenode recovery, capacity planning, and slots configuration.
  • Created HBase tables to store variable data formats of PII data coming from different portfolios.
  • Implemented a script to transmit sysprin information from Oracle to Hbase using Sqoop.
  • Implemented best income logic using Pig scripts and UDFs.
  • Implemented test scripts to support test driven development and continuous integration.
  • Worked on tuning the performance Pig queries.
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
  • Responsible to manage data coming from different sources.
  • Involved in loading data from UNIX file system to HDFS.
  • Load and transform large sets of structured, semi structured and unstructured data
  • Cluster coordination services through Zookeeper.
  • Experience in managing and reviewing Hadoop log files.
  • Installed Oozie workflow engine to run multiple Hive and pig jobs.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.

Environment: Hadoop, HDFS, Pig, Hive, Sqoop, HBase, Shell Scripting, Ubuntu, Linux Red Hat.

Confidential

Java Developer

Responsibilities:
  • Involved in Design, Development and Support of the application used AGILE methodology and participated in SCRUM meetings.
  • Extensively used My Eclipse as an IDE for building, developing and integrating the application.
  • Extensively used Rally'sAgile Management tool (Rally Dev).
  • Provided JUnit test cases for the application to support the Test Driven Development (TDD).
  • Manipulated DB2 for data retrieving and storing using ORM.
  • Developed Web Service client interface with JAX-RPC from WSDL files for invoking the methods using SOAP.
  • Extensively worked on SOA andWeb Services in Axis 2.0 to get the data from third party systems.
  • Provided SQL scripts and PL/SQL stored procedures for querying the database.
  • Provide Maven, MS build tool for building and deploying the application.
  • Building and Deployed the application in Web logic Application Server.
  • Created system architecture and design using the UML Analysis Model and Design Model.
  • Developed Servlets and a JSP for performing CRUD operations on domain specific entities.
  • Developed Data Access Layer using Hibernate and DAO Design Pattern.
  • Extensively used Spring IOC architectural model to inject objects based on the selection of components like setter injection and Interface injection to manage the object references.
  • Involved in the development of the application based on backend Spring MVC architecture.
  • Utilized SpringMVC framework to implement design patterns like IOC (Dependency Injection), Spring DAO (Data access objects), Data Transfer objects, Business objects, ORM Mappings.
  • Design to reuse Spring framework starting from user submitting the HTTP Servlet request from JSP and Dispatcher Servlet passing the request to Controller to service layer and delegating the request to DAO layer for via Facade using Business Delegator Design Pattern.
  • Used the Spring DAO to handle exception for database transaction like open connections, no result, connection aborted, closing the connections etc.
  • Used Design Patterns like value object, session Facade and Factory.
  • Developed the presentation Tier using JSP, XHTML, and HTML.
  • Third party credit card information accessed via SOAP Web-Services.
  • Check-in and Checkout of application is achieved using CVS.

Environment: OOAD, Java 1.6, J2EE, HTML, XHTML, CSS, JavaScript, AJAX, JQuery, Spring 3.0, Maven2, DAO, DTO, Web logic, JPA, JAX-WS, SOAP UI, SVN, JBOSS, Spring MVC, JUnit 4.

We'd love your feedback!