We provide IT Staff Augmentation Services!

Sr. Hadoop Developer Resume

3.00/5 (Submit Your Rating)

Basking Ridge, NJ

SUMMARY

  • 8+ years of professional IT experience in analysis, design, architecture, development, testing and implementation of Big Data, Data Warehousing, Business Intelligence and deploying various applications on Object Oriented Programming.
  • Expertise in Hadoop, Map Reduce, Spark - Scala, YARN, Spark Stream, Hive, Pig, HBase, Kafka, Cassandra, Oracle.
  • Hands on experience wif major components in Hadoop Ecosystem including Hive, HBase, Pig, Sqoop, Flume, Avro, Oozie, Zookeeper and MapReduce frameworks, Cassandra, Apache NIFI.
  • Worked of Hadoop Architecture and various components such as HDFS, JobTracker, TaskTracker, NameNode, DataNode and MapReduce concepts.
  • Worked on NoSQL databases like MongoDB, HBase, and Cassandra.
  • Well versed in configuring and administrating the Hadoop Cluster using major Hadoop Distributions like Cloudera and Hortonworks.
  • Installed, Configured and Maintained Apache Hadoop clusters for application development and Hadoop tools like Hive, Pig, HBase, Oozie, Flume, Zookeeper and Sqoop.
  • Strong noledge on setting up and managing the batch scheduler on Oozie.
  • Experience in analyzing data using HiveQL, PigLatin and custom MapReduce programs in Java, Python.
  • Hands on experience in writing Pig UDFs, Hive UDFs and UDAF’s in the analysis of data.
  • Experienced on loading large sets of structured, semi structured and unstructured data and also performed importing and exporting data into HDFS and Hive using Sqoop.
  • Hands on experience in working wif Flume to load the log data from multiple sources directly into HDFS.
  • Hands on experience in Apache NIFI to automate the flow of data between systems.
  • Experience in implementing Data Warehousing/ETL solutions for different domains like Health, Banking, Financial, and insurance verticals.
  • Design, investigation and implementation of public facing websites on Amazon Web Services AWS.
  • Have an experience on writing python scripts to update content in the database and manipulate files.
  • Designed and developed the monitoring system used for AWS Elastic Environments, including monitoring individual instance health, overall environment health, and managing individual component failure cases e.g. missing ELB, misconfigured Auto scaling Group, misconfigured EC2 Security Group, etc.
  • Experience in Integration of Amazon Web Services AWS wif other applications infrastructure.
  • Hands on experience in writing MR jobs for cleansing the data and to copy it to AWS cluster form our cluster.
  • Performed basic GIS operations including creating/editing geographic data sets, using geo-processing tools.
  • Proficient in developing web based applications and client server distributed architecture applications in Java/J2EE technologies using Object Oriented Methodology.
  • Strong Knowledge on full Software Development life cycle-Software analysis, design, architecture, development and maintenance.
  • Worked on relative ease wif different working strategies like Agile, Waterfall, Scrum and Test Driven Development (TDD) methodologies.
  • Excellent experience in designing and developing Enterprise Applications for J2EE platform using Servlets, JSP, JDBC, Struts, Spring, Hibernate and Web services.
  • Expertise in client-side design and validations using HTML, DHTML, AJAX, CSS, Java Script, JSP, Struts Tag Library.
  • Hands on experience in working on XML Suite of technologies (XML, XSD, DTD, XML Schema, DOM).
  • Expertise in developing web services wif XML based protocols such as SOAP and WSDL.
  • Very strong Business Modeling skills using Rational Unified Process, OOAD and UML.
  • Hands on experience in Application Development using Java, Hadoop, RDBMS and Linux shell scripting.
  • Good working noledge on relational databases like MySQL, Oracle and NoSQL databases like HBase.
  • Extensively used different IDEs like Eclipse, Net Beans and RAD.
  • Experience wif design Patterns like MVC, Singleton, Factory, Proxy, DAO, Abstract, Prototype and Adaptor.
  • Strong noledge on to prepare scripts to ensure proper data access, manipulation and reporting functions wif Object Oriented Programming languages.

TECHNICAL SKILLS

Big Data Technologies: HDFS, Hive, Map Reduce, Pig, Sqoop, Oozie, Zookeeper, YARN, Avro, Spark, Impala, Flume, Ambari, Kafka

Scripting Languages: Shell, Python, Perl, Scala, R Programming

Programming Languages: Java, C++, C, SQL, PL/SQL. Pig Latin, Hive QL

Front End Technologies: HTML, XHTML, CSS, XML, JavaScript, AJAX, Servlets, JSP.

Java Frameworks: MVC, Apache Struts2.0, Spring and Hibernate.

Web Services: SOAP (JAX-WS), WSDL, SOA, Restful (JAX-RS), JMS, AWS.

Application Servers: Apache Tomcat, WebLogic Server, WebSphere, JBoss.

Databases: Oracle 11g, MySQL, MS SQL Server, IBM DB2.

NoSQL Databases: HBase, MongoDB, Cassandra.

IDE: Eclipse, NetBeans, RAD, JBuilder.

RDBMS: Oracle 9i, Oracle 10g, MS Access, MS SQL Server, IBM DB2, PL/SQL.

Operating Systems: Linux, UNIX, MAC, Windows

Networks: HTTP, HTTPS, FTP, UDP, TCP/TP, SNMP, SMTP, POP3.

PROFESSIONAL EXPERIENCE

Confidential, Basking Ridge, NJ

Sr. Hadoop Developer

Responsibilities:

  • Gatheird the business requirements from the Business Partners and Subject Matter Experts.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Experienced on loading and transforming of large sets of structured, semi structured and unstructured data.
  • Experienced in Big Data technologies such as Hadoop, Cassandra, Presto, Spark, Flume, Storm, AWS and SQL.
  • Created Hive tables to store the processed results in a tabular format.
  • Involved in HDFS maintenance and loading of structured and unstructured data.
  • Developed several Map Reduce Programs for data preprocessing
  • Wrote Hive queries for data analysis to meet the business requirements.
  • Prepared System Design document wif all functional implementations.
  • Involved in Data model sessions to develop models for HIVE tables.
  • Understanding the existing Enterprise data warehouse set up and provided design and architecture suggestion converting to Hadoop using MapReduce, HIVE, SQOOP and Pig Latin.
  • Converting ETL logic to Hadoop mappings.
  • Developed a process for the Batch ingestion of CSV Files, Sqoop from different sources and also generating views on the data source using Shell Scripting and Python.
  • Knowledge on handling Hive queries using Spark SQL that integrate wif Spark environment implemented in Scala.
  • Daily status check for Oozie workflow and monitor Cloudera manager and check data node status to ensure nodes are up and running.
  • Experience wif leveraging Hadoop ecosystem components including Pig and Hive for Data Analysis, Sqoop for Data Migration, Oozie for Data Scheduling, and HBase as a NoSQL data store.
  • Worked wif cloud services like Amazon Web Services (AWS) and involved in ETL, Data Integration and Migration.
  • Extensive hands on experience in Hadoop file system commands for file handling operations.
  • Worked on Sequence files, Map side joins, bucketing, partitioning for hive performance enhancement and storage improvement.Loading from disparate data sets and pre-processed using Hive and Pig.
  • Extracted the data from Teradata into HDFS using Sqoop.
  • Performed analysis of data sets and uncovered insights and maintained security and data privacy.
  • Implemented data flow scripts using UNIX / hive / pig scripting.
  • Developed a script in Scala to read all the Parquet Tables in a Database and parse them as JSON files, another script to parse them as structured tables in Hive.
  • Loaded data in elastic search from Data Lake using Spark/Hive for Kibana.
  • Developer for full text search platform using NoSQL Elastic search engine, allowing for much faster, more scalable and more intuitive user searches for our database of spas worldwide.
  • Deployed Hadoop cluster of Hortonworks Distribution and installed ecosystem components: HDFS, YARN, Zookeeper, HBase, Hive, MapReduce, Pig, Kafka, Storm and Spark in Linux servers using Ambari.
  • Used spatial analysis to identify potential sites for macro dams, schools, health facilities and roads.
  • Introduced seven digit codes for Eritrean towns for database analysis and created a standardized dictionary solving misunderstandings of multiple spellings for town names.
  • Configured Zookeeper for Cluster co-ordination services.
  • Designed a standard map template allowing for a simplified process to produce standardized maps still currently in use.
  • Provided solutions to the customer to streamline data to work across multiple software platforms.
  • Analyzed and refined efficient search query algorithms to implement business requirements.
  • Efficiently handled periodic exporting of SQL data into Elastic search.
  • Proposed best practices/standards and worked wif the System Analyst and development manager on a day-to-day basis.

Environment: Hadoop, MapReduce, Hive, HDFS, PIG, Spark, Oozie, Scala, Python, Flume, Yarn, Zookeeper, Kafka, Sqoop, MapReduce, HBase, Java and Linux Shell scripting, Talend, Teradata, Unix/Linux.

Confidential, Phoenix, AZ

Sr. Hadoop Developer

Responsibilities:

  • Installed and configured Hadoop Map Reduce, HDFS and developed multiple Map Reduce jobs in Java for data cleaning and preprocessing
  • Collaborate wif subject matter experts, various stakeholders and fellow developers to design, develop, implement and support data analytics.
  • Importing and exporting data into HDFS and Hive using Sqoop, Spark Core and Spark SQL
  • Created Hive tables, and loading and analyzing data using hive queries
  • Having exposure to Teradata for processing the huge data
  • Worked on debugging, performance tuning of Hive & Pig Jobs.
  • Designed and implemented pig UDFs for evaluation, filtering, loading and storing of data
  • Worked on Performance Tuning ofHadoopjobs by applying techniques such as MapSide Joins, Partitioning, Bucketing and using different file formats such as SequenceFile, RCFile, ORCFile
  • Defined job work flows as per their dependencies in OOZIE.
  • Used JAVA, J2EE application development skills wif Object Oriented Analysis and extensively involved throughout Software Development Life Cycle (SDLC).
  • Designed and implemented the MongoDB schema.
  • Installed Apache Tez, a programing framework which is built on YARN in increase performance.
  • Experience on deployment of Apache Tez on top of YARN.
  • Wrote services to store and retrieve user data from the MongoDB for the application on devices.
  • Used Mongoose API in order to access the MongoDB from NodeJS.
  • Used different Scala Collection API’s to process the data in Spark Streaming.
  • Developed Spark code using Scala for faster data processing using RDD's and Data frame API.
  • Used Python and Django to interface wif the jQuery UI and manage the storage and deletion of content.
  • Proactively monitored systems and services, architecture design and implementation of Hadoop deployment, configuration management, backup, and disaster recovery systems and procedures.
  • Installed and configure Zookeeper service for coordinating configuration-related information of all the nodes in the cluster to manage it efficiently.
  • Used Flume to collect, aggregate, and store the web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
  • Load and transform large sets of structured, semi structured and unstructured data
  • Supported Map Reduce Programs those are running on the cluster
  • Involved in loading data from UNIX file system to HDFS, configuring Hive and writing Hive UDFs
  • Processing the streamed data using real time messaging systems Kafka and Strom
  • Utilized Java and ORACLE from day to day to debug and fix issues wif client processes
  • Managed and reviewed log files
  • Having good amount of noledge on Cassandra, HBase.
  • Implemented partitioning, dynamic partitions and buckets in HIVE

Environment: Hadoop, JDK1.6, Map Reduce, HDFS, Hive, Strom, Mongo DB, Zookeeper, Scala, Python, Kafka, Cassandra, Pig, Spark core, Spark SQL, Sqoop, Yarn, HTML, XML, SQL, J2EE, Eclipse, RC, ORC, Flume, Thrift, Oozie, HBase.

Confidential, Plano, TX

Sr. Hadoop/Spark Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop.
  • Responsible for managing and scheduling Jobs on a Hadoop cluster.
  • Loading data from UNIX file system to HDFS and vice versa.
  • Improving the performance and optimization of existing algorithms in Hadoop using Spark context, Spark-SQL and Spark YARN.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
  • Worked wif Apache Spark for large data processing integrated wif functional programming language Scala.
  • Developed POC using Scala, Spark SQL and MLlib libraries along wif Kafka and other tools as per requirement tan deployed on the Yarn cluster.
  • Extract Real time feed using Kafka and Spark Streaming and convert it to RDD and process data in the form of Data Frame and save the data as Parquet format in HDFS.
  • Implemented Data Ingestion in real time processing using Kafka.
  • Worked on Oozie workflow engine for job scheduling.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster processing of data
  • Configured Spark Streaming to receive real time data and store the stream data to HDFS.
  • Developed Spark scripts by using Scala shell commands as per the requirement
  • Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
  • Documented the requirements including the available code which should be implemented using Spark, Hive, HDFS and SOLR.
  • Utilized Python scripting to automate GIS workflows.
  • Implemented High Availability and automatic failover infrastructure to overcome single point of failure for Name node utilizing Zookeeper services.
  • Tested Apache TEZ, an extensible framework for building high performance batch and interactive data processing applications, on Pig and Hive jobs.
  • Used Kafka Streams to Configure Spark streaming to get information and tan store it in HDFS.
  • Developed multiple Kafka Producers and Consumers as per the software requirement specifications.
  • Extract Real time feed using Kafka and Spark Streaming and convert it to RDD and process data in the form of Data Frame and save the data as Parquet format in HDFS.
  • Real time streaming the data using Spark wif Kafka.
  • Responsible for creating Hive tables and working on them using Hive QL.
  • Implementing various Hive UDF’s as per business requirements.
  • Involved in Data Visualization using Tableau for Reporting from Hive Tables.
  • Developed Python Mapper and Reducer scripts and implemented them using Hadoop Streaming.
  • Developed multiple Map Reduce jobs in java for data cleaning and preprocessing.
  • Optimized Map Reduce Jobs to use HDFS efficiently by using various compression mechanisms.
  • Responsible for writing Hive queries for data analysis to meet the business requirements.
  • Customized Apache Solr to handle fallback searching and provide custom functions.
  • Responsible for setup and benchmarking of Hadoop/HBase clusters.

Environment: Hadoop, HDFS, HBase, Sqoop, Hive, Zookeeper, Yarn, Map Reduce, Oozie, Spark- Streaming/SQL, Scala, Kafka, Solr, Sbt, Java, Python, Ubuntu/Cent OS, MySQL, Linux, GitHub, Maven, Jenkins.

Confidential

Java/ Hadoop Developer

Responsibilities:

  • Involved in Requirements Analysis, and design an Object-oriented domain model.
  • Involvement in the detailed Documentation, written functional specifications of the module.
  • Involved in development of Application wif Java and J2EE technologies.
  • Develop and maintain elaborate services based architecture utilizing open source technologies like Hibernate, ORM and Spring Framework.
  • Developed server-side services using Java multithreading, Struts MVC, Java, EJB, and spring, Web Services (SOAP, WSDL, and AXIS).
  • Responsible for developing DAO layer using Spring MVC and configuration XML’s for Hibernate and to also manage CRUD operations (insert, update, and delete).
  • Designing, Development and Implementation of JSPs in Presentation layer for Submission, Application, and implementation.
  • Development of JavaScript for client end data entry validations and Front-End Validation.
  • Deployed Web, presentation and business components on Apache Tomcat Application Server.
  • Developed PL/SQL procedures for different use case scenarios.
  • Involvement in post-production support, Testing and used JUNIT for unit testing of the module.

Environment: Java/J2EE, JSP, XML, Spring Framework, Hibernate, Eclipse(IDE), Java Script, Ant, SQL, PL/SQL, Oracle, Windows, UNIX, Soap, Jasper reports.

Confidential 

Software Developer

Responsibilities:

  • Involved in Design, Development and Support phases of Software Development Life Cycle (SDLC).
  • Reviewed the functional, design, source code and test specifications.
  • Involved in developing the complete front end development using Java Script and CSS.
  • Author for Functional, Design and Test Specifications.
  • Developed web components using JSP, Servlets and JDBC.
  • Designed tables and indexes.
  • Designed, Implemented, Tested and Deployed Enterprise Java Beans both Session and Entity using WebLogic as Application Server.
  • Developed stored procedures, packages and database triggers to enforce data integrity. Performed data analysis and created crystal reports for user requirements.
  • Implemented the presentation layer wif HTML, XHTML and JavaScript.
  • Implemented Backend, Configuration DAO, XML generation modules of DIS.
  • Analyzed, designed and developed the component.
  • Used JDBC for database access.
  • Used Spring Framework for developing the application and used JDBC to map to Oracle database.
  • Used Data Transfer Object (DTO) design patterns.
  • Unit testing and rigorous integration testing of the whole application.
  • Written and executed the Test Scripts using JUNIT.
  • Actively involved in system testing.
  • Developed XML parsing tool for regression testing.
  • Prepared the Installation, Customer guide and Configuration document which were delivered to the customer along wif the product.

Environment: Java, JavaScript, HTML, CSS, JDK 1.5.1, JDBC, Oracle10g, XML, XSL, Solaris and UML.

We'd love your feedback!