We provide IT Staff Augmentation Services!

Hadoop Developer Resume

2.00/5 (Submit Your Rating)

San Jose, CA

SUMMARY:

  • 7+ years of IT professional experience with emphasis on Design, Development, Implementation, Testing, Maintenance and Deployment of Software Applications using Java, J2EE and Big Data technologies.
  • Experience in Big Data technologies and Hadoop ecosystem components like Spark, HDFS, MapReduce, Pig, Hive, YARN, Sqoop, Flume, Kafka and NoSQL systems like HBase, Cassandra.
  • AWS Certified Developer.
  • Expertise in architecting real time streaming applications and batch style large scale distributed computing applications using tools like Spark Streaming, Kafka, Hive, Impala etc.
  • Hands - on experience in installing, configuring and monitoring HDFS clusters (on premise & cloud AWS).
  • Experience in the IT industry as a Ab-Initio Developer.
  • Developed prototype for parsing high volume XML files using Hadoop and storing output in HDFS for Ab Initio.
  • Implemented detailed systems and services monitoring using Nagios, Zabbix, & AWS Cloud Watch.
  • In depth understanding of Map Reduce and AWS cloud concepts and its critical role in data analysis of huge and complex datasets.
  • Created custom Database Encryption Connectors that could be plugged into Sqoop to be able to encrypt the data while importing to HDFS/Hive.
  • Integrated Maven with Jenkins for the builds as the Continuous Integration process .
  • Experience in fine-tuning and troubleshooting Spark Applications, Hive queries.
  • Extensive hands on experience in writing complex MapReduce jobs, Pig Scripts and HiveQL scripts.
  • Experience using various Hadoop Distributions (Cloudera, Hortonworks, Amazon AWS) to fully implement and leverage new Hadoop features.
  • Experience in Continuous Integration and Deployments (CI/CD) using build tools like Jenkins, TeamCity, MAVEN, and ANT. Wrote scripts to automate Build.
  • Build, manage, and continuously improved the build infrastructure for global software development engineering teams including implementation of build scripts, continuous integration infrastructure and deployment tools.
  • Experience in using Amazon Cloud services like S3, EMR etc.
  • Experience in Apache Flume and Kafka for collection, aggregation and moving huge chunks of data from various sources such as web server, telnet sources etc.
  • Worked on Java HBase API for ingesting processed data to HBase tables.
  • Extensive knowledge and experience in using Apache Storm, Spark Streaming, Apache Spark, Apache Nifi, Kafka and Flume in creating data streaming solutions.
  • Experienced in working with Machine learning libraries (spark-MLlib) and implementing ML algorithms for clustering, regression filtering and dimensional reduction.
  • Extensive understanding of Partitions and Bucketing concepts in Hive.
  • Created few Hive UDF's to perform some complex business specific transformations and rules.
  • Expert knowledge over J2EE Design Patterns like MVC Architecture, Session Facade, Front Controller and Data Access Objects for building J2EE Applications.
  • Experienced in using Agile methodologies including extreme programming, Scrum Process and Test-Driven Development (TDD).
  • Use of Ansible for environment automation, configuration management and provisioning Setting up playbooks to deploy, manage, test and configure software onto the hosts.
  • Integrate Data Meer with tools and cloud-based platforms within the big data ecosystem (e.g. Spark, Tez, Azure, HDI, Amazon EMR, RedShift, Google Data Proc)
  • Intensive work experience in developing enterprise solutions using Java, J2EE, Struts, Servlets, Hibernate, JavaBeans, JDBC, JSP, JSF, JSTL, MVC, Spring, Custom Tag Libraries, JNDI, AJAX, SQL, JavaScript, AngularJS and XML.
  • Conversant with web application Servers like Tomcat, WebSphere, WebLogic and Jboss servers.
  • Experience in development of logging standards and mechanism based on Log4j.
  • Experience in writing ANT and Maven scripts to build and deploy Java applications.

TECHNICAL SKILLS:

Big Data: Hadoop, HDFS, Spark, Hive, Impala, Kafka, Hue, MapReduce, YARN, Pig, Sqoop, HBase, Couchbase, Cassandra, Oozie, Storm, Flume, Talend, AWS, Hortonworks and Cloudera clusters

Language: Scala, Java, C, UNIX Shell Scripting, AngularJS, PL/SQL, Python

Java/J2EE: J2EE, JSF, EJB, HTML, XHTML, AngularJS, Servlets, JSP, CSS, XML, Ajax, Java script, SOAP, RESTful

Open source framework and web development: Struts, Spring, Hibernate, JavaScript, AJAX, Dojo, JQuery, Ehcache, Log4j, Ant, JBoss, Web services, SOA, SOAP, REST, WSDL and UDDI

Portals/Application servers: WebLogic, WebSphere Application server, WebSphere Portal server, JBOSS

Operating system: Windows, AIX, UNIX, Linux.

ETL Tools: Ab Initio GDE Version 3.0.2.2, Co>Operating system 3.1.6.1, EME, Data Profiler, Familiarity with Ab Initio ACE Application Configuration Environment and BRE Business Rule Engine.

Configuration Mgmt.: CMVC, ClearCase, Clearquest, PVCS, CVS, Nagios, Puppet, Ansible.

Development Tools: Eclipse, Visual Studio, Net Beans, Rational Application Developer, WSAD, Junit.

Databases: Couchbase, Cassandra, HBase, Oracle 10g, MySQL, Teradata SQL.

Software Engineering: UML 2.0, Rational Rose, Design Patterns (MVC, DAO etc.).

PROFESSIONAL EXPERIENCE:

Hadoop Developer

Confidential, San Jose, CA

Responsibilities:

  • Worked on importing and exporting data from Teradata, MySQL into HIVE using Sqoop for visualization, analysis and to generate reports.
  • Loaded CSV files containing user event information into Hive External tables on daily basis.
  • Created Spark applications and used Data frames and Spark-SQL API primarily for performing event enrichment and performing lookups with other enterprise data sources.
  • Build a real time streaming pipeline by using Kafka integration with Storm and Spark Streaming.
  • Perform ETL on different formats of data like JSON, CSV files and converted them to parquet while loading to final tables. Ran ad-hoc querying using Hive and Impala.
  • Extensively perform complex data transformations in Spark using Scala language.
  • Worked extensively with importing metadata into Hive using Scala and migrated existing tables and applications to work on Hive and AWS cloud.
  • Involved in converting Hive/SQL queries into Spark transformations using Scala.
  • Connect Tableau and Squirrel SQL clients to Spark-SQL (Spark thrift server) via data source and run the queries.
  • Extensively worked on Informatica IDE/IDQ.
  • Involved in massive data profiling using IDQ (Analyst Tool) prior to data staging.
  • Used IDQ to profile the project source data, define or confirm the definition of the metadata, cleanse and accuracy check the project data, check for duplicate or redundant records, and provide information on how to proceed with ETL processes .
  • Extensively worked with the Ab Initio Enterprise Meta Environment EME to obtain the initial setup variables and maintaining version control during the development effort.
  • Designed Ab Initio graphs that would harness Teradata capabilities ELT and balance resources between database and Ab Initio.
  • Worked with Machine learning libraries (Spark MLlib) and done clustering, regression, filtering and dimensional reduction using implemented ML algorithms.
  • Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scala (Prototype).
  • Used Impala as the primary analytical tool for allowing visualization servers to connect and perform reporting on top Hadoop directly.
  • Used Machine learning libraries as Spark MLlib and implemented ML algorithms for clustering, regression filtering and dimensional reduction process with data scientists.
  • Installed and configured monitoring tools Nagios for critical applications.
  • Wrote spark jobs for Data clustering and data processing using Spark-MLlib and cluster algorithms as per functional requirements.
  • Working on Oozie workflow engine to run multiple Hive-QL jobs and on schedulers.
  • Done various compressions and file formats like Parquet, Snappy, Gzip, Bzip2, Avro, Text.
  • Implemented test scripts to support test driven development and continuous integration.
  • Used Zookeeper to provide coordination services to the cluster.
  • Used Impala and Tableau to create various reporting dashboards.

Environment: HDFS, Hadoop 2.x, Pig, Hive, Sqoop, Flume, Spark, MapReduce, Scala, Oozie, YARN, Tableau, Spark-SQL, Spark-MLlib, Impala, Nagios, UNIX Shell Scripting, Zookeeper, Kafka, Agile Methodology, Cloudera 5.9, SBT.s

Confidential, Westerville

Hadoop Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop.
  • Implemented Installation and configuration of multi-node cluster on Cloud using Amazon Web Services (AWS) on EC2.
  • Worked on importing and exporting data from DB2 into Hive using Sqoop.
  • Involved in importing the real-time data to Hadoop using Kafka and implemented the Oozie job for daily imports.
  • Developed Hive scripts to perform data transformation and etl processes.
  • Developed MapReduce (YARN) programs to cleanse the data in HDFS obtained from assorted data sources to make it suitable for ingestion into Hive schema for analysis.
  • Designed and developed MapReduce jobs to process data coming in different file formats like XML, CSV, JSON.
  • Extensively worked on Informatica IDE/IDQ.
  • Involved in massive data profiling using IDQ (Analyst Tool) prior to data staging.
  • Used IDQ’s standardized plans for addresses and names clean ups.
  • Worked on IDQ file configuration at user’s machines and resolved the issues.
  • Used IDQ to complete initial data profiling and removing duplicate data.
  • And extensively worked on IDQ admin tasks and worked as both IDQ Admin and IDQ developer.
  • Transferring data between MySQL and Hadoop Distributed File System (HDFS) using Sqoop with connectors.
  • Creating and populating Hive tables and writing hive queries for data analysis to meet the business requirements.
  • Running MapReduce jobs to access HBase data from application using Java Client APIs.
  • Developed Pig Latin scripts and HQL queries for the analysis of Structured, Semi-Structured and Unstructured data.
  • Automating the extraction, processing and analysis of data jobs using Oozie. Used SVN for version control.
  • Deploying and managing applications in Datacentre, Virtual environment and Azure platform as well.
  • Experience in large-scale streaming data analytics using Storm.
  • Created working POC’s using Spark 1.1.0 streaming for real time stream processing of continuous stream of large data sets.
  • Experience using Apache Nifi to track dataflows from beginning to end.
  • Involved building and managing NoSQL Database like Hbase or Cassandra .
  • Integrating Cassandra with Storm for real time user attributes look up.
  • Worked in Spark to read the data from Hive and write it to Cassandra using Java.
  • Worked on Azure for highly available customer facing B2B and B2C applications.
  • Developed custom aggregate functions using Spark nifi and performed interactive querying.
  • Developed Spark scripts by using Scala and Python Shell commands as per the requirement.
  • Installed and configured Hadoop Scala, HDFS (AWS cloud) and responsible for Developing multiple Spark jobs in Scala, R scripts & CLI for data managing and pre-processing events.
  • Implemented Machine learning libraries, spark-MLlib algorithms for clustering, filtering, regression and dimensional reduction.
  • Worked on data warehouse product Amazon Redshift which is a part of the AWS (Amazon Web Services)
  • Developed Spark and Spark SQL scripts to migrate data from RDBMS into AWS-RedShift.
  • Worked with Data Analysts for the integration of external data sources with Hadoop using Denodo.
  • Developed ETL Scripts for Data acquisition and Transformation using Talend.

Environment: Hadoop 2.x, Hive, HQL, HDFS, MapReduce, Spark 1.1.0, Scala, Sqoop, Storm, Kafka, Nifi, Flume, Oozie, HBase, AWS-RedShift, Python, Java, Maven, Eclipse, Putty, AWS-EC2, Cassandra, Talend, Azure, Hortonworks, Puppet, Ansible

Confidential, Dallas, TX

Hadoop Developer

Responsibilities:

  • Analysing the functional specifications provided by the client and developing detailed solution design document with the Architect and the team.
  • Used Hadoop architecture with Map Reduce functionality and its ecosystem to solve the customer requirements using Cloudera Distribution for Hadoop (CDH).
  • Developed multiple Map Reduce jobs in Java for complex business requirements including data cleansing and pre-processing.
  • Created Data Lake which serves as a base layer to store and do analytics on data flowing from multiple sources into Hadoop Platform.
  • Involved in creating Hive tables, loading the data and writing hive queries which will run internally in map reduce.
  • Optimized Hive tables using optimization techniques like partitions and bucketing to provide better performance with HiveQL queries.
  • Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
  • Scripted complex HiveQL queries on Hive tables for analytical functions by implementing Hive Generic UDFs.
  • Wrote customized User Defined Function’s (UDF) to process further in Java and Perl to ease the processing in Pig.
  • Involved in building applications using Maven and integrated them by using continuous Integration servers like Jenkins to build jobs.
  • Worked on Import and export data from Legacy Databases RDBMS into HDFS and Hive using Sqoop.
  • Worked extensively with importing metadata into Hive and migrated existing tables and applications to work on Hive and AWS cloud.
  • Designing NoSQL schemas in HBase and Cassandra.
  • Involved in Agile methodologies, daily Scrum meetings, Sprint planning.

Environment: Hadoop 1.x, HDFS, MapReduce, Hive, Pig, HBase, Tez, Sqoop, Oozie, Maven, Shell Scripting, Teradata, CDH3, Cloudera Manager.

Confidential, Pennsylvania

Hadoop Developer

Responsibilities:

  • Implementing project using Agile SCRUM methodology, involved in daily stand up meetings and sprint showcase and sprint retrospective.
  • Developed the web tier using JSP, Spring MVC. Used Spring Framework for the Implementation of the Application.
  • Integrated Spring Dependency Injection (IOC) among different layers of an application.
  • Used Hibernate for object Relational Mapping and used JPA for annotations.
  • Implemented REST web services using Apache-CXF framework.
  • Involved in creating various Data Access Objects (DAO) for addition, modification and deletion of records using various specification files.
  • Developed presentation layer using HTML, JSP, Ajax, CSS and JQuery.
  • Deployed the Application in WebSphere server.
  • Designed and developed persistence layer using spring JDBC template.
  • Involved in Unit Testing of various modules in generating the Test Cases.
  • Used SVN and GitHub as version control tool.
  • Used Maven for build and management. Extensively involved in Test-Driven Development (TDD).
  • Converted the HTML Pages to JSF Tag Specific Pages. Developed JSPs and managed beans using JSF.

Environment: Spring framework, Spring MVC, Spring JDBC, Hibernate, J2EE, JSP, Ajax, XML, Log4j Maven, JavaScript, HTML, CSS, JQuery, PL/SQL, SVN, GitHub, WebSphere, Agile, JAX-WS, Apache-CXF.

Confidential

Java Developer

Responsibilities:

  • Implemented the project according to the Software Development Life Cycle (SDLC)
  • Implemented JDBC for mapping an object-oriented domain model to a traditional relational database
  • Created Stored Procedures to manipulate the database and to apply the business logic according to the user’s specifications
  • Developed the Generic Classes, which includes the frequently used functionality, so that it can be reusable
  • Exception Management mechanism using Exception Handling Application Blocks to handle the exceptions
  • Designed and developed user interfaces using JSP, Java script and HTML
  • Involved in Database design and developing SQL Queries, stored procedures on MySQL
  • Used CVS for maintaining the Source Code
  • Logging was done through log4j

Environment: JAVA, Java Script, HTML, log4j, JDBC Drivers, Soap Web Services, Unix, Shell scripting, SQL Server

We'd love your feedback!