We provide IT Staff Augmentation Services!

Sr. Hadoop Developer Resume

0/5 (Submit Your Rating)

Chicago, IL

SUMMARY

  • Around 8 Years of IT experience in Application Development domain of Java and Big Data.
  • Experience in Hadoop And it’s Eco System; HDFS, Map Reduce, Apache Pig, Hive, HBase, Oozie, Scala, Spark, Flume, Kafka, Storm And Sqoop.
  • Experienced in the Hadoop ecosystem components like Hadoop Map Reduce, Cloudera, Hortonworks, HBase, Oozie, Hive, Sqoop, Pig, Tez, Flume, Kafka, Storm, Spark, Scala, MongoDB, Couchbase and Cassandra.
  • Experience in analyzing data using HIVEQL and Pig Latin and custom Map Reduce programs in Java and scala.
  • Good knowledge on building Apache spark applications using Scala.
  • Having a Good exposure on Big Data technologies andHadoopecosystem, In - depth understanding of Map Reduce and theHadoopInfrastructure.
  • Excellent knowledge onHadoopArchitecture and ecosystems such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
  • Strong experience in architecting real time streaming applications and batch style large scale distributed computing applications using tools like Spark Streaming, Spark SQL, mlib, Kafka, Flume, Map reduce, Hive etc.
  • Cassandra developer: Set-up configured and optimized the Cassandra cluster. Developed real-time java based application to work along with the Cassandra database.
  • Proficient in Cassandra Data Modeling and Analysis and CQL (Cassandra Query Language). Have 2 years of profound experience in Cassandra database
  • Hadoop Administrator and developer: Set-up configured and monitored Hadoop cluster. Executed performance tuning of cluster.
  • Skilled in managing and reviewing Hadoop log files.
  • Expert in importing and exporting data into HDFS and Hive using Sqoop.
  • Experienced in loading data to Hive partitions and creating buckets in Hive
  • Experienced in configuring Flume to stream data into HDFS.
  • Experienced in real-time Big Data solutions using Hbase, handling billions of records.
  • Extensive experience in working with structured/semi-structured and Unstructured data by implementing complex map reduce programs using design patterns.
  • Familiarity with Hadoop architecture and its components like HDFS, Map Reduce, Job Tracker, Task Tracker, Name Node and Data Node.
  • Experienced in Application Development using Java, Hadoop, RDBMS and Linux shell scripting and performance tuning.
  • Expertise in writing ETL Jobs for analyzing data using Pig.
  • Experienced querying in Impala.
  • Experience in data-warehousing with ETL tool Oracle Warehouse Builder (OWB).
  • Familiarity with distributed coordination system Zookeeper.
  • Experienced in implementing unified data platforms using Kafka producers/ consumers, implement pre-processing using storm topologies.
  • Experienced in working with Solr indexing and querying.
  • Excellent knowledge in Java and SQL in application development and deployment.
  • In-depth understanding of Data Structures and Algorithms and Optimization.
  • Have a very good understanding and worked with relational databases like MySQL, Oracle and NoSQL databases like Hbase, Mongo DB, Couchbase and Cassandra.
  • Well versed with databaseslike MS SQL Servers 2012 and 2008, Oracle 11g/10g/9i, MySQL.
  • Passionate towards working inHadoopand Big Data Technologies, data science, machine learning in Spark, Big Data Processing, Analytics and Visualization.
  • Versatile experience in utilizing Java tools in business, web and client server environments including Java platform, JSP, Servlets, Java beans and JDBC.
  • Expertise in developing the presentation layer components like HTML, CSS, JavaScript, JQuery, XML, JSON, AJAX and D3.
  • Experienced in source control repositories viz. SVN, GitHub.
  • Good Knowledge of analyzing data in HBase using Hive and Pig.
  • Experienced in detailed system design using use case analysis, functional analysis, modelling program with class and sequence, activity and state diagrams using UML.
  • Worked with Data-Warehouse Architecture and Designing Star Schema, Snow flake Schema, Fact and Dimensional Tables, Physical and Logical Data Modeling.
  • Designed Mapping documents for Big Data Application.
  • Experienced in Agile and SCRUM.

TECHNICAL SKILLS

Big Data Eco System: HDFS, Map Reduce, Hive, Pig, HBase, Spark, Spark Streaming, Spark SQL, Kafka, Cloudera CDH4, CDH5, Hortonworks, Hadoop Streaming, Zookeeper, Oozie, Sqoop, Flume, Impala, Solr, Tez and Ranger.

No SQL: HBase, MongoDB, Couchbase, Neo4j, Cassandra

Languages: Java/ J2EE, SQL, Shell Scripting, C/C++, Python, Scala

Web Technologies: HTML, JavaScript, CSS, XML, Servlets, SOAP, Amazon AWS, Google App Engine

Web/ Application Server: Apache Tomcat Server, LDAP, JBOSS, IIS

Operating system: Windows, Macintosh, Linux and Unix

Frameworks: Springs, MVC, Hibernate, Swings

DBMS / RDBMS: Oracle 11g/10g/9i, SQL Server 2012/2008, MySQL

IDE: Eclipse, Microsoft Visual Studio (2008,2012), NetBeans, Spring Tool Suits

Version Control: SVN, CVS and Rational Clear Case Remote Client, GitHub, Visual Studio

Tools: FileZilla, Putty, TOAD SQL Client, MySQL Workbench, ETL, DWH, JUnit, SQL Oracle Developer, WinScp, Tahiti, Cygwin

PROFESSIONAL EXPERIENCE

Confidential, Chicago IL

Sr. Hadoop Developer

Responsibilities:

  • Moved Relational Database data using Sqoop as ETL tool into the Data Lake, Hive Dynamic partition tables.
  • Used Recursive queries using several levels of joins in Hive to transform the SQL Stored Procedures, to make them work in the data lake.
  • Optimizing the Hive queries using Partitioning, Bucketing techniques with ACID properties for controlling the data distribution.
  • Running Hive queries through different engines like Spark, MapReduce and TEZ.
  • Worked with NoSQL database Hive, Hbase to create tables and store data.
  • Worked on custom Pig Loaders and storage classes to work with variety of data formats such as JSON and XML file formats.
  • Experienced in connecting web services in JSON with the databases and extracting data using java.
  • Used Hortonworks distribution of hadoop.
  • Experienced in Using Pig as ETL tool to do Transformations, even joins and some pre-aggregations before storing the data onto HDFS.
  • Used Kerberos Authentication system through Ranger.
  • Used Distributed copy for intra-cluster data transfer.
  • Used Oozie workflow engine and coordinators to manage timed interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive and Sqoop.
  • Developed workflow in Oozie to automate the tasks of loading the data into HDFS and processing with Sqoop and Hive.
  • Hive performance tuning for the best efficient results.
  • Used Map Reduce programs using Chained Mappers to create data pipeline.
  • Implemented Various Optimization techniques in hive on TEZ for effective result sets.
  • Implemented Optimized join base by joining different data sets to get top claims based on state using Map Reduce.
  • Create a complete processing engine, based on Hortonworks’ distribution, enhanced to performance.
  • Experienced in Monitoring Cluster health using Ambari.
  • Developed data pipeline using Flume, Sqoop, Pig and Java MapReduce to ingest behavioral data into HDFS for analysis.
  • Consolidating employee data from Clients, Consultants and Employees around the globe, into Data Lake and visualize the resource trend report through qlickview.
  • Responsible for importing large sets of data from MySQL and load into HDFS using Sqoop on regular basis.
  • Created customized BI tool for manager team that perform Query analytics using HiveQL.
  • Created Partitions, Buckets based on State to further process using Bucket based Hive joins.
  • Saved the storage space on HDFS by using compression techniques like LZO, GZip, ZLib and SNAPPY.
  • Compressed the data to fit into the infrastructure with minimal hardware requirements.
  • Created Hive Generic (One to Many and Many to One) UDF's, UDAF's, UDTF's in java to process business logic that varies based on conventions.
  • Worked on shell scripting in Linux and the Cluster. Used shell scripts to run hive queries from beeline.
  • Created SQL stored procedure’s programmability in HiveQL to run analytics on the data imported.
  • Used various transformations like Filter, Expression, Sequence Generator, Update Strategy, Joiner, Stored Procedure, and Union to develop robust mappings in the Informatica Designer.
  • Worked on improving performance of Hive and Pig queries.
  • Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
  • Designing conceptual model with Spark for performance optimization.
  • Implemented custom codes for map reduce partitioner and custom writables.
  • Implemented test scripts to support test driven development and continuous integration.
  • Trained and Mentored analyst / test team for writing, running and validating Sqoop scripts and Hive Queries.
  • Worked on Hive for exposing data for further analysis and for generating transforming files from different analytical formats to text files.
  • Implemented a script to transmit sys print information from MySQL to HBase using Sqoop.
  • Implemented Spark using Scala and SparkSQL for faster testing and processing of data.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs and Scala.
  • Written programs in scala that runs in spark and worked on Hue interface for querying the data.
  • Developed Spark jobs using Scala in test environment for faster data processing and used Spark SQL for querying.
  • Worked in Linux/Unix Environment.
  • Worked in Agile development team environment, in Sprints with daily scrum meetings.
  • Technical documentation around the whole environment.

Environment: Hadoop, HDFS, HBase, MapReduce, Java, REST, Spark, Hive, Beeline, TEZ, Pig, Sqoop, Flume, Oozie, Hue, Zookeeper, Ambari, Java, Scala, SQL, ETL, DWH, Hortonworks, HUE, Ranger, MySQL.

Confidential, Atlanta GA

Sr. Hadoop Developer

Responsibilities:

  • Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data.
  • Developed data pipeline using Flume, Sqoop, Pig and Java MapReduce to ingest behavioral data into HDFS for analysis.
  • Responsible for importing log files from various sources into HDFS using Flume.
  • Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
  • Extracted files from MongoDB through Sqoop and placed in HDFS and processed.
  • Created customized BI tool for manager team that perform Query analytics using HiveQL.
  • Created Partitions, Buckets based on State to further process using Bucket based Hive joins.
  • Estimated the hardware requirements for Name Node and Data Nodes & planning the cluster.
  • Created Hive Generic UDF's, UDAF's, UDTF's in java to process business logic that varies based on policy.
  • Moved Relational Database data using Sqoop into Hive Dynamic partition tables using staging tables.
  • Consolidating customer data from Lending, Insurance, Trading and Billing systems into data warehouse and mart subsequently for business intelligence reporting.
  • Optimizing the Hive queries using Partitioning and Bucketing techniques, for controlling the data distribution.
  • Worked with NoSQL database Hbase to create tables and store data.
  • Proficient in querying Hbase using Impala.
  • Worked on custom Pig Loaders and storage classes to work with variety of data formats such as JSON and XML file formats.
  • Used Pig as ETL tool to do Transformations, even joins and some pre-aggregations before storing the data onto HDFS.
  • Design technical solution for real-time analytics using Kafka and Hbase.
  • Used Pig as ETL tool to do transformations, event joins, filter and some pre-aggregations.
  • Collaborated with Business users for requirement gathering for building Tableau reports per business needs.
  • Experience in Upgrading Apache Ambari, CDH and HDP Cluster.
  • Configured and Maintained different topologies in storm cluster and deployed them on regular basis.
  • Imported structured data, tables into Hbase.
  • Experienced with different kind of compression techniques like LZO, GZip, and Snappy.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
  • Created Data Pipeline of Map Reduce programs using Chained Mappers.
  • Configuring Spark Streaming to receive real time data from the Kafka and Store the stream data to HDFS.
  • Implemented Spark RDD transformations, actions to migrate Map reduce algorithms.
  • Set-up configured and optimized the Cassandra cluster. Developed real-time java based application to work along with the Cassandra database.
  • Worked around the Cassandra database and proficient in in Cassandra Data Modeling and Analysis and CQL (Cassandra Query Language).
  • Implemented Optimized join base by joining different data sets to get top claims based on state using Map Reduce.
  • Converting queries to Spark SQL and using parquet file as storage format.
  • Developed analytical component using Scala, Spark and SparkStream.
  • Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scala.
  • Written spark programs in Scala and ran spark jobs on YARN.
  • Designed and Implemented Solr Search using the big data pipeline.
  • Assembled Hive and Hbase with Solr to build a full pipeline for data analysis
  • Experienced in sync up Solr with HBase to compute indexed views for data exploration.
  • Implemented map reduce programs to perform joins on the Map side using Distributed Cache in Java. Developed Unit test cases using Junit, Easy Mock and MRUnit testing frameworks.
  • Used in depth features of Tableau like Data Blending from multiple data sources to attain data analysis.
  • Experience in upgrading hadoop cluster hbase/zookeeper from CDH3 to CDH4.
  • Used Maven extensively for building MapReduce jar files and deployed it to Amazon Web Services (AWS) using EC2 virtual Servers in the cloud.
  • Used Maven extensively for building MapReduce jar files and deployed it to Amazon Web Services (AWS) using EC2 virtual Servers in the cloud and Experience in build scripts to do continuous integrations systems. Had an exposure to Amazon Web Services - AWS cloud computing (EMR, EC2 and S3 services).
  • Setup Amazon web services (AWS) to check whether Hadoop is a feasible solution or not.
  • Worked with BI teams in generating the reports and designing ETL workflows on Tableau.
  • Create a complete processing engine, based on Cloudera's distribution, enhanced to performance.
  • Experienced in Monitoring Cluster using Cloudera manager.

Environment: Hadoop, HDFS, HBase, MapReduce, Java, JDK 1.5, J2EE 1.4, Struts 1.3, Spark, Hive, Pig, Sqoop, Flume, Impala, Oozie, Hue, Solr, Zookeeper, Kafka, AVRO Files, SQL, ETL, DWH, Cloudera Manager, MySQL, Scala, MongoDB.

Confidential, Irvine, CA

Hadoop Developer

Responsibilities:

  • Installed and configured Hadoop and Hadoop stack on a 16 node cluster.
  • Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables.
  • Involved in data ingestion into HDFS using Sqoop from variety of sources using the connectors like JDBC and import parameters.
  • Analyze large and critical datasets of Global Risk Investment and Treasury Technology (GRITT) Domain using Cloudera, HDFS, Hbase, MapReduce, Hive, Hive UDF, Pig, Sqoop, Zookeeper, & Mahout.
  • Designed and implemented MapReduce-based large-scale parallel relation-learning system.
  • Worked with NoSQL databases like Hbase in creating Hbase tables to load large sets of semi structured data coming from various sources.
  • Extracted files from CouchDB through Sqoop and placed in HDFS and processed.
  • Involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs
  • Exported the data from Avro files and indexed the documents in sequence file format.
  • Implemented various performance optimizations like using distributed cache for small datasets, Partitioning, Bucketing in hive, using Compression Codecs where ever necessary
  • Implemented test scripts to support test driven development and continuous integration
  • Install, configure, and operate data integration and analytic tools i.e. Informatica, Chorus, SQLFire, & Gem Fire XD for business needs.
  • Implementing MapReduce programs to analyse large datasets in data warehouse for business intelligence purpose
  • Involved in designing and developing non-trivial ETL processes within Hadoop using tools likePig, Sqoop, Flume, and Oozie
  • Excellent experience in ETL analysis, designing, developing, testing and implementing ETL processes including performance tuning and query optimizing of database
  • Experience in using Pentaho Data Integration tool for data integration, OLAP analysis and ETL process.
  • Develop scripts to automate routine DBA tasks (i.e. refresh, backups, vacuuming, etc.)
  • Installed and configured Hive and also wrote Hive UDF’s that helped spot market trends.
  • Used Hadoop streaming to process terabytes data in XML format.
  • Involved in loading data from UNIX file system to HDFS.
  • Implemented Fair schedulers on the Job tracker with appropriate parameters to share the resources of the Cluster for the Map Reduce jobs given by the users.
  • Designed and developed a decision tree application using Neo4J graph database for fraud visuals
  • Worked on generating reports Neo4J graph database.
  • Classification of key words into categories using a Neo4J graph database and make a recommendation.
  • Involved in creating Hive tables, loading the data using it and in writing Hive queries to analyze the data.
  • Worked on installing cluster, commissioning & decommissioning of Data nodes, Name node recovery, capacity planning, JVM tuning, map and slots configuration.
  • Extensively involved in installation and configuration of Cloudera distribution of Hadoop, its Name node, Secondary Name node, Job tracker, Task trackers and Data nodes.
  • Installed and configured Hadoop ecosystem like HBase, Flume, Pig and Sqoop.
  • Worked on analyzing Hadoop stack and different big data analytic tools including Pig and Hive, HBase database and Sqoop.
  • Gained very good business knowledge on different category of products and designs within.
  • Created various Documents such as Source-To-Target Data mapping Document, Unit Test Cases and Data Migration Document.
  • Created mappings using the transformations like Source Qualifier, Aggregator, Expression, Lookup, Router, Normalizer, Filter, Update Strategy and Joiner transformations.
  • Designed high level ETL architecture for overall data transfer from the OLTP to OLAP.

Environment: CDH4 with Hadoop 1.x, HDFS, Pig, Cloudera, Hive, Hbase, Zookeeper, MapReduce, Java, Sqoop, Oozie, ETL, DWH, Neo4j, Ambari, Linux, UNIX Shell Scripting and Big Data.

Confidential, Atlanta, GA

Java Developer

Responsibilities:

  • Involved in Analysis, Design and Development of the project.
  • Designed and developed web-based software using Java Server Faces (JSF) framework, Spring MVC Framework, and Spring Web Flow.
  • Developed user interface using JSP, HTML, XHTML and Java Script to simplify the complexities of the application.
  • Used Ajax for intensive user operations and client-side validations.
  • Developed the code using object oriented programming concepts.
  • Developed application service components and configured beans using Spring IoC, creation of Hibernate mapping files and generation of database schema.
  • Used Web Services for creating rate summary and used WSDL and SOAP messages for getting insurance plans from different module and used XML parsers for data retrieval.
  • Used JUnit for testing the web application.
  • Used JAXM for making distributed software applications communicate via SOAP and XML
  • Used DB2 as backend database.
  • Used SQL statements and procedures to fetch the data from the DB2 database.
  • Involved in writing Spring Configuration XML, file that contains declarations and business classes are wired-up to the frontend-managed beans using Spring IOC pattern.
  • Involved in creating various Data Access Objects (DAO) for addition, modification and deletion of records using various specification files.
  • Developed Ant Scripts for the build process and deployed in IBM WebSphere.
  • Implemented Log4J for Logging Errors, debugging and tracking using loggers, appenders’ components.
  • Implemented Business processes such as user authentication, Transfer of Service using Session EJBs.
  • Involved in the Bug fixing of various applications reported by the testing teams in the application during the integration and used Bugzilla for the bug tracking.
  • Used Tortoise CVS as version control across common source code used by developers.
  • Deployed the applications on IBM Web Sphere Application Server.

Environment: JDK1.5, J2EE, Spring 2.0, Servlets, JSP, MAVEN, Hibernate 3.0, JSON, XML, Swing, ANT, JSF, JMS, EJB, WSDL, DB2, JUnit, Apache, Tomcat, CVS, Log4J, RAD 7.0, Eclipse, WAS.

Confidential

Software Developer

Responsibilities:

  • Developed, enhanced and tested of the Web Methods flow services and Java services.
  • Used WebServices for interaction between various components and created SOAP envelopes.
  • Created web service XML connectors for use within a flow service
  • Created a front-end application using JSPs and Spring MVC for registering a new transaction and configured it to connect to database
  • Used Spring Security for Authorization of users and implemented Spring WebServices.
  • Developed and configured Microsoft SQL Server 2008 tables including Sequences, Functions, Procedures and Table constraints.
  • Created standalone Java application to read data from several XLS files and insert data into the Database as needed by the Testing team.
  • Ant Build tool configuration for automation of building processes for all types of environment - Test, QA, and Production.
  • Developed and provided support to many components of this application from end-to-end, i.e. Front-end (View) to Web Methods and Database.
  • Provided solutions for bug fixes in this application.
  • Developed queries & triggers in the database.
  • Used Tortoise SVN as a version-controlling tool for managing the module developments.

Environment: Java 1.5, spring 2.5, Spring WebServices, Spring Security, Web logic, Hibernate 3.0, SQL Server 2008, Eclipse, XML, JSON, Apache Ant.

Confidential

Java Developer

Responsibilities:

  • Designed, developed and enhanced back-end system using J2EE and Database technologies for Subscriber Sales and Marketing operations
  • Experienced in using Spring Beans, DAO, Controller and Service layers for fast and efficient processing for retail store sales processing
  • Worked with Spring Dependency Injection (DI), Annotations, AOP and advanced multi-threaded application request processing
  • Reviewed and managed Entity Relationship Diagram, Entity Beans, and Spring declarative transaction models
  • Wrote and optimized SQL queries using Oracle Hints and created supportive stored procedures using PLSQL, Triggers and Cursors
  • Designed and developed WebServices using Axis, and associated business modules integration for B2B partner application interfaces
  • Implemented persistence application layer using Hibernate, HQL, Criteria, Annotations, and Object Relational Model (ORM) model.
  • Worked on cross-domain application technology based on Web 2.0 and current Bell platform RESTful API
  • Worked in advanced requirements analysis (BAs), rapid application creation and deployment teams.
  • Analyzed complex data call flows, transaction management and application data traffic models.

Environment: Java, JSP, spring, Hibernate, JavaScript, HQL Struts, Servlets, Axis, Eclipse, Ant, JDBC, WebServices, Web logic 8.1, Oracle 9i, SQL Plus.

We'd love your feedback!