We provide IT Staff Augmentation Services!

Bigdata/hadoop Developer. Resume

4.00/5 (Submit Your Rating)

Washington, DC

PROFESSIONAL SUMMARY:

  • Having 8+ years of overall experience in building and developing Hadoop MapReduce solutions.
  • Strong experience with Big Data and Hadoop technologies with excellent knowledge of Hadoop ecosystem: Hive, Spark, Sqoop, Impala, Pig, HBase, Kafka, Flume, Storm, Zookeeper, Oozie.
  • Designed, configured and deployed Amazon Web Services (AWS) for a multitude of applications utilizing the AWS stack (Including EC2, Route53, S3, RDS, CloudFormation, Cloud Watch, SQS, IAM), focusing on high - availability, fault tolerance, and auto-scaling.
  • Worked on HDFS, namenode, job tracker, data node, TASKTRACKER and the Map-Reduce concepts.
  • Experience in ingestion, storage, querying, processing and analysis of Big Data with hands-on experience in Big Data including Apache Spark, Spark SQL, and Spark Streaming.
  • Worked with Spark engine to process large-scale data and experience to create Spark RDD.
  • Excellent experience on Installation, Configuration, and Administration of Hadoop cluster of major Hadoop distributions such as Cloudera Enterprise (CDH3 and CDH4) and Hortonworks Data Platform (HDP1 and HDP2).
  • Diverse experience utilizing Java tools in business, Web, and client-server environments including Java Platform, J2EE, EJB, JSP, Java Servlets, Struts, and Java Database Connectivity (JDBC) technologies.
  • Experienced in using Zookeeper and Oozie Operational Services for coordinating the cluster and scheduling workflows.
  • Expertise in creating several UDF's, UDAF, UDTF using Java and developing Machine learning algorithms using Mahout for clustering and data mining
  • Experienced in Installation, configuration, and administration of Hadoop Cluster.
  • Experienced in supporting Apache and Tomcat applications running on Linux and Unix servers and support of applications running on Linux machines
  • Knowledge of developing Spark Streaming jobs by using RDDs and leverage Spark-Shell.
  • Expertise in Talend Big data tool involved in architectural designing and development of ingestion and extraction job in Big Data and Spark Streaming.
  • Having experience on RDD architecture and implementing Spark operations on RDD and optimizing transformations and actions in Spark.
  • Experienced working with Hive/HQL to query data from Hive tables in HDFS and successfully loaded files to Hive and HDFS from Oracle and SQL Server using Sqoop.
  • Expertise in using NoSQL database HBase, Cassandra for storing large tables by bringing data to HBase using Pig and Sqoop
  • Experience in generating On-demand and Scheduled Reports for business analysis or Management decision using SQL Server Reporting Services ( SSRS ).
  • Worked on designing complex reports including subreports and formulas with complex logic using SQL Server Reporting Services ( SSRS ).
  • Experience in Hadoop administration activities such as installation and configuration of clusters using Apache and Cloudera.
  • Hands-on Apache Spark jobs using Scala in a test environment for faster data processing and used SparkSQL for querying.
  • Good in analyzing data using HiveQL, Pig Latin and custom MapReduce program in Java.
  • Good in Hive and Impala queries to load and processing data in Hadoop Filesystem (HFS).
  • Good understanding of NoSQL Databases and hands-on work experience in writing applications on NoSQL databases like Cassandra and MongoDB.
  • Experience in building high performance and scalable solutions using various Hadoop ecosystem tools like Pig, Hive, Sqoop, Spark, Solr, and Kafka.
  • Hands on experience in installing, configuring, and using Hadoop ecosystem components like Hadoop MapReduce, HDFS, HBase, Oozie, Hive, Sqoop, Pig, Zookeeper, and Apache Storm.
  • Good Knowledge of HDFS high availability (HA) and different demons of Hadoop clusters which include Resource Manager, Node Manager, Name Node and Data Node
  • Experience with all stages of the SDLC and Agile Development model right from the requirement gathering to Deployment and production support.
  • Hands-on Agile (Scrum), Waterfall model along with automation and enterprise tools like Jenkins, Chef, JIRA, Confluence to develop projects and version control, Git.

TECHNICAL SKILLS:

Bigdata Ecosystem: Hadoop, HDFS, YARN, MapReduce, Hive, Pig, Impala, Sqoop, Flume, Spark, Kafka, Storm, Zookeeper, and Oozie

Languages: HTML5, DHTML, WSDL, CSS3, C, C++, XML, R/R Studio, SAS, Schemas, JSON, Ajax, Java, Scala, Python, Shell Scripting

NO SQL Databases: Cassandra, HBase, MongoDB, MariaDB

Business Intelligence Tools: QlikView, Amazon Redshift, or Azure Data Warehouse

Development Tools: Microsoft SQL Studio, IntelliJ, Eclipse, NetBeans.

Development Methodologies: Agile/Scrum, UML, Design Patterns, Waterfall.

Build Tools: Toad, SQL Loader, Maven, ANT.

Databases: Microsoft SQL Server 2008,2010/2012, MySQL 4.x/5.x, Oracle 11g, 12c, DB2, Teradata, Netezza

Operating Systems: Windows, UNIX, LINUX.

PROFESSIONAL EXPERIENCE:

Confidential, Washington, DC

Bigdata/Hadoop Developer

Responsibilities:

  • Responsible for installation and configuration of Hive, Pig, HBase and Sqoop on the Hadoop cluster and created hive tables to store the processed results in a tabular format.
  • Configured Spark Streaming to receive real-time data from the Apache Kafka and store the stream data to HDFS using Scala.
  • Developed the Sqoop scripts to make the interaction between Hive and vertical Database.
  • Processed data into HDFS by developing solutions and analyzed the data using Map Reduce PIG, and Hive to produce summary results from Hadoop to downstream systems.
  • Build servers using AWS: Importing volumes, launching EC2, creating security groups, auto-scaling, load balancers, Route 53, SES and SNS in the defined virtual private connection.
  • Written Map Reduce code to process and parsing the data from various sources and storing parsed data into HBase and Hive using HBase-Hive Integration.
  • Streamed AWS log group into Lambda function to create service now incident.
  • Involved in loading and transforming large sets of Structured, Semi-Structured and Unstructured data and analyzed them by running Hive queries and Pig scripts.
  • Created Managed tables and External tables in Hive and loaded data from HDFS.
  • Developed Spark code by using Scala and Spark-SQL for faster processing and testing and performed complex HiveQL queries on Hive tables.
  • Scheduled several times based Oozie workflow by developing Python scripts.
  • Developed Pig Latin scripts using operators such as LOAD, STORE, DUMP, FILTER, DISTINCT, FOREACH, GENERATE, GROUP, COGROUP, ORDER, LIMIT, UNION, SPLIT to extract data from data files to load into HDFS.
  • Exporting the data using Sqoop to RDBMS servers and processed that data for ETL operations.
  • Worked on S3 buckets on AWS to store Cloud Formation Templates and worked on AWS to create EC2 instances.
  • Designing ETL Data Pipeline flow to ingest the data from RDBMS source to Hadoop using a shell script, Sqoop, package, and MySQL.
  • End-to-end architecture and implementation of client-server systems using Scala, Akka, Java, JavaScript and related, Linux
  • Optimized the Hive tables using optimization techniques like partitions and bucketing to provide better.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
  • Implementing Hadoop with the AWS EC2 system using a few instances in gathering and analyzing data log files.
  • Involved in Spark and Spark Streaming creating RDD's, applying operations -Transformation and Actions.
  • Created partitioned tables and loaded data using both static partition and dynamic partition method.
  • Developed custom Apache Spark programs in Scala to analyze and transform unstructured data.
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and Extracted the data from Oracle into HDFS using Sqoop
  • Using Kafka on publish-subscribe messaging as a distributed commit log, have experienced in its fast, scalable and durability.
  • Test Driven Development (TDD) process and extensive experience with Agile and SCRUM programming methodology.
  • Implemented POC to migrate Map Reduce jobs into Spark RDD transformations using SCALA
  • Scheduled map reduces jobs in a production environment using Oozie scheduler.
  • Involved in Cluster maintenance, Cluster Monitoring, and Troubleshooting, Manage and review data backups and log files.
  • Designed and implemented map reduce jobs to support distributed processing using java, Hive and Apache Pig
  • Analyzing Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, and Sqoop.
  • Improved the Performance by tuning of HIVE and map reduce.
  • Research, evaluate and utilize modern technologies/tools/frameworks around Hadoop ecosystem.

Environment: HDFS, Map Reduce, Hive, Sqoop, Pig, Flume, Vertica, Oozie Scheduler, Java, Shell Scripts, Teradata, Oracle, HBase, MongoDB, Cassandra, Cloudera, AWS, JavaScript, JSP, Kafka, Spark, Scala and ETL, Python.

Confidential - Collegeville, PA

Hadoop/Spark Developer.

Responsibilities:

  • Hands on experience in Spark and Spark Streaming creating RDD & applying operations transformations and Actions.
  • Developed Spark applications using Scala for easy Hadoop transitions.
  • Used Spark and Spark-SQL to read the parquet data and create the tables in hive using the Scala API.
  • Performed advanced procedures like text analytics and processing, using the in-memory computing capabilities of Spark using Scala.
  • Developed Spark code using Scala and Spark-SQL for faster processing and testing.
  • Implemented Spark sample programs in python using PySpark.
  • Analyzed the SQL scripts and designed the solution to implement using PySpark.
  • Developed PySpark code to mimic the transformations performed in the on-premise environment.
  • Used Spark-Streaming APIs to perform necessary transformations and actions on the fly for building the common learner data model which gets the data from Kafka in near real time.
  • Responsible for loading Data pipelines from web servers and Teradata using Sqoop with Kafka and Spark Streaming API.
  • Developed Kafka producer and consumers, Cassandra clients and Spark along with components on HDFS, Hive.
  • Populated HDFS and HBase with huge amounts of data using Apache Kafka.
  • Used Kafka to ingest data into Spark engine.
  • Configured deployed and maintained multi-node Dev and Test Kafka Clusters.
  • Managing and scheduling Spark Jobs on a Hadoop Cluster using Oozie.
  • Experienced with different scripting languages like Python and shell scripts.
  • Developed various Python scripts to find vulnerabilities with SQL Queries by doing SQL injection, permission checks, and performance analysis.
  • Tested Apache TEZ, an extensible framework for building high-performance batch and interactive data processing applications, on Pig and Hive jobs.
  • Designed and implemented Incremental Imports into Hive tables and writing Hive queries to run on TEZ.
  • Experienced data pipelines using Kafka and Akka for handling large terabytes of data.
  • Written shell scripts that run multiple Hive jobs which helps to automate different Hive tables incrementally which are used to generate different reports using Tableau for the Business use.
  • Experienced in Apache Spark for implementing advanced procedures like text analytics and processing using the in-memory computing capabilities written in Scala.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDD, Scala, and Python.
  • Worked on SparkSQL, created Data frames by loading data from Hive tables and created prep data and stored in AWS S3.
  • Using Spark-Streaming APIs to perform transformations and actions on the fly for building the common learner data model which gets the data from Kafka in near real-time and Persists into Cassandra.
  • Involvement in creating custom UDFs for Pig and Hive to consolidate strategies and usefulness of Python into Pig Latin and HQL (HiveQL).
  • Implemented Hortonworks Nifi (HDP 2.4) and recommended a solution to inject data from multiple data sources to HDFS and Hive using Nifi.
  • Developed various data loading strategies and performed various transformations for analyzing the datasets by using Hortonworks Distribution for Hadoop ecosystem.
  • Ingested data from RDBMS and performed data transformations, and then export the transformed data to Cassandra as per the business requirement and used Cassandra through Java services.
  • Experience in NoSQL Column-Oriented Databases like Cassandra and its Integration with Hadoop cluster.
  • Wrote ETL jobs to read from web APIs using REST and HTTP calls and loaded into HDFS using java and Talend.
  • Along with the Infrastructure team, involved in the design and developed Kafka and Storm based data pipeline.
  • Involved in loading and transforming large Datasets from relational databases into HDFS and vice-versa using Sqoop imports and export.
  • Written shell scripts that run multiple Hive jobs which helps to automate different Hive tables incrementally which are used to generate different reports using Tableau for the Business use.
  • Used Hibernate ORM framework with Spring framework for data persistence and transaction management.

Environment: Hadoop, Hive, MapReduce, Sqoop, Kafka, Spark, Yarn, Pig, Cassandra, Oozie, shell Scripting, Scala, Maven, Java, JUnit, agile methodologies, NIFI, MySQL, Tableau, AWS, EC2, S3, Hortonworks, Power BI.

Confidential - Herndon, VA

Hadoop/Spark Developer.

Responsibilities:

  • Optimizing of existing algorithms in Hadoop using Spark context, Spark-SQL, Data Frames and Pair RDD's.
  • Developed Spark scripts by using Java, and Python shell commands as per the requirement.
  • Involved with ingesting data received from various relational database providers, on HDFS for analysis and other big data operations.
  • Experienced in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
  • Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
  • Worked on Spark SQL and Data frames for faster execution of Hive queries using Spark SQL Context.
  • Performed analysis of implementing Spark using Scala.
  • Responsible for creating, modifying topics (Kafka Queues) as and when required with varying configurations involving replication factors and partitions.
  • Extracted files from MongoDB through Sqoop and placed in HDFS and processed.
  • Created and imported various collections, documents into MongoDB and performed various actions like query, project, aggregation, sort, and limit.
  • Experience with creating a script for data Modeling and data import and export. Extensive experience in deploying, managing and developing MongoDB clusters.
  • Extracted files from MongoDB through Sqoop and placed in HDFS and processed.
  • Experience in migrating HiveQL into Impala to minimize query response time.
  • Creating Hive tables to import large data sets from various relational databases using Sqoop and export the analyzed data back for visualization and report generation by the BI team.
  • Involved in creating Shell scripts to simplify the execution of all other scripts (Pig, Hive, Sqoop, Impala, and MapReduce) and move the data inside and outside of HDFS.
  • Collecting data from various Flume agents that are imported into various servers using Multihop Flow.
  • Developed Scala scripts, UDFs using Data frames/SQL/Datasets and RDD/MapReduce in Spark 1.6 for Data Aggregation, queries and writing data back into OLTP system through Sqoop.
  • Developed Pig Scripts, Pig UDF and Hive Scripts, Hive UDFs to analyze HDFS data.
  • Maintained the cluster securely using Kerberos and making the cluster up and running all the times.
  • Implemented optimization and performance testing and tuning of Hive and Pig.
  • Developed a data pipeline using Kafka to store data into HDFS.
  • Worked on reading multiple data formats on HDFS using Scala
  • Written shell scripts and Python scripts for automation of job.

Environment: Cloudera, HDFS, Hive, HQL scripts, MapReduce, Java, HBase, Pig, Sqoop, Kafka, Impala, Shell Scripts, Python Scripts, Spark, Scala, Oozie.

Confidential

Hadoop developer

Responsibilities:

  • Collected and aggregated large amounts of weblog data from different sources such as web servers, mobile and network devices using Apache Flume and stored the data into HDFS for analysis.
  • Collecting data from various Flume agents that are imported on various servers using Multi-hop flow.
  • Ingest real-time and near-real-time ( NRT ) streaming data into HDFS using Flume.
  • Extensively involved in Installation and configuration of Cloudera distribution Hadoop, Name node, Secondary Name Node, Job Tracker, Task Trackers and Data Nodes.
  • Developed MapReduce programs in Java and Sqoop the data from ORACLE database.
  • Responsible for building Scalable distributed data solutions using Hadoop. Written various Hive and Pig scripts.
  • Used various HBase commands and generated different Datasets as per requirements and provided access to the data when required using grant and revoke.
  • Created HBase tables to store variable data formats of input data coming from different portfolios.
  • Worked on HBase for support enterprise production and loading data into HBase using SQOOP.
  • Installed Oozie workflow engine to run multiple Hive and Pig jobs which run independently with time and data availability.
  • Experienced in analyzing Cassandra database and compare it with other open-source NoSQL databases to find which one of them better suites the current requirements.
  • Implemented Storm builder topologies to perform cleansing operations before moving data into Cassandra.
  • Performed Sqooping for various file transfers through the HBase tables for processing of data to several NoSQL DBs- Cassandra, MongoDB.
  • Expertise in understanding Partitions, Bucketing concepts in Hive.
  • Experience working with Apache SOLR for indexing and querying.
  • Created custom SOLR Query segments to optimize ideal search matching.
  • Used Oozie Scheduler system to automate the pipeline workflow and orchestrate the MapReduce Jobs that extract the data in a timely manner. Responsible for loading data from UNIX file system to HDFS.
  • Developed suit of Unit Test Cases for Mapper, Reducer and Driver classes using MR Testing library.
  • Analyzed the weblog data using the HiveQL, integrated Oozie with the rest of the Hadoop stack
  • Utilized cluster co-ordination services through Zookeeper.
  • Worked on the Ingestion of Files into HDFS from remote systems using MFT.
  • Got good experience with various NoSQL databases and Comprehensive knowledge in process improvement, normalization/de-normalization, data extraction, data cleansing, and data manipulation.
  • Developed Pig scripts to convert the data from Text file to Avro format.
  • Created Partitioned Hive tables and worked on them using HiveQL .
  • Developed Shell scripts to automate routine DBA tasks.
  • Used Maven extensively for building jar files of MapReduce programs and deployed to Cluster.
  • Responsible for cluster maintenance, adding and removing cluster nodes, cluster monitoring and troubleshooting, managing and reviewing data backups and Hadoop log files.

Environment: HDFS, MapReduce, Pig, Hive, Oozie, Sqoop, Flume, HBase, Java, Maven, Avro, Cloudera, Eclipse and Shell Scripting.

Confidential

Java Developer.

Responsibilities:

  • Installation, Configuration & Upgrade of Solaris and Linux operating system.
  • Actively participated in requirements gathering, analysis, design, and testing phases
  • Designed use case diagrams, class diagrams, and sequence diagrams as a part of Design Phase
  • Developed the entire application implementing MVC Architecture integrating JSF with Hibernate and Spring frameworks.
  • Developed the Enterprise Java Beans (Stateless Session beans) to handle different transactions such as online funds transfer, bill payments to the service providers.
  • Implemented Service Oriented Architecture (SOA) using JMS for sending and receiving messages while creating web services
  • Developed XML documents and generated XSL files for Payment Transaction and Reserve Transaction systems.
  • Developed SQL queries and stored procedures.
  • Developed Web Services for data transfer from client to server and vice versa using Apache Axis, SOAP, and WSDL.
  • Used JUnit Framework for the unit testing of all the java classes.
  • Implemented various J2EE Design patterns like Singleton, Service Locator, DAO, and SOA.
  • Worked on AJAX to develop an interactive Web Application and JavaScript for Data Validations.
  • Developed the application under JEE architecture, developed Designed dynamic and browser compatible user interfaces using JSP, Custom Tags, HTML, CSS, and JavaScript.
  • Deployed & maintained the JSP, Servlets components on Web logic 8.0
  • Developed Application Servers persistence layer using, JDBC, SQL, Hibernate.
  • Used JDBC to connect the web applications to Data Bases.
  • Implemented Test First unit testing framework driven using Junit.
  • Developed and utilized J2EE Services and JMS components for messaging communication in Web Logic.
  • Configured development environment using Web logic application server for developer's integration testing
  • Created unit test cases and worked with the QA team in designing the test cases and testing of the migrated code.

Environment: Java/J2EE, SQL, Oracle 10g, JSP 2.0, EJB, AJAX, Java Script, Web Logic 8.0, HTML, JDBC 3.0, XML, JMS, log4j, Junit, Servlets, MVC, My Eclipse.

Confidential

Java Developer.

Responsibilities:

  • Involved in various stages of the SDLC phases of Analysis, design, coding, testing and deploying the application.
  • Followed Waterfall methodology for developing the application.
  • Worked on Hibernate framework to persist the data.
  • Used SQL statements and procedures to manipulate & retrieve the data from the Oracle 8i database.
  • Developed the major front-end modules for the application using HTML, JSP using Struts 1.1 Framework.
  • Involved in creating web pages using HTML, CSS, JavaScript, and jQuery.
  • Used JavaScript for Client-side validations and validation framework for server-side validations.
  • Used AJAX for the Web page rendering.
  • Co-ordination with the onsite team for developing, testing and production issues.
  • Used WebSphere application server for deployment of the application.
  • Used CVS for Version control system.
  • Fixed the bugs in the testing phase.
  • Developed POJO classes for defining the variables in the object class.
  • Writing the Unit Testing of the components using JUnit.
  • Used log4j for logging messages to the log files.
  • Developed ANT scripts and developed builds using Apache ANT.

Environment: Eclipse, HTML, JavaScript, Core Java, JUnit, JSP, jQuery, Servlets, JDBC, SQL, Oracle 8i, AJAX, CVS and WebSphere Application Server, Log4j

We'd love your feedback!