Hadoop Developer Resume
Buffalo New, YorK
SUMMARY
- 7+ years of IT industry experience and 5+ years of expertise in Big Data platforms, architectures and systems
- Experienced in CDH, HDP and CDP ecosystem including HDFS, MapReduce, Yarn, Hive, Sqoop, HBase, Cassandra, Spark, Oozie, Zookeeper, Kafka, etc.
- Implemented, managed and maintained on - premise and cloud-based Hadoop clusters
- Configured Spark ingestion and streaming from Kafka with Flume and stored processed data to HDFS
- Leveraged Spark transformations, actions, UDFs, DAG and data lineage to achieve business goals
- Converted HQL queries into Spark transformations using Spark RDDs and DataFrames in Scala
- Tuned Spark performance, troubleshoot Spark issues and managed Spark sesources
- Used Spark DataFrames API over Cloudera platform to perform analytics on Hive data
- Experienced in importing and exporting the data using Sqoop to and from HDFS to RDBMS
- Experienced in RDBMSs like MySQL, MS SQL Server, Oracle Server and Teradata
- Good understanding and knowledge in NoSQL databases like MongoDB, HBase and Cassandra
- Experienced in workflow scheduling and locking tools/services like Oozie and Zookeeper
- Practiced ETL methods in enterprise-wide solutions, data warehousing, reporting and data analysis
- Used apache NiFi to copy the data from local file system to HDFS
- Written Hive UDFs as required and executed complex HiveQL queries to extract data from Hive tables
- Leveraged partitioning, bucketing and both managed and external tables in Hive for performance optimization
- Used Pig scripts for transformations, event joins, filters and pre-aggregations before storing the data onto HDFS
- Experienced in working with AWS using EMR, EC2, RDS, RedShift for computation and S3 as storage mechanism
- Good Knowledge in Unix Shell scripting for automating deployments and other routine tasks
- Experienced in using IDEs like Eclipse, NetBeans, IntelliJ IDEA and Spring framework
- Leveraged Jira and Rally for management and tracking and Git, BitBuket and and SVN for source code management
- Worked in all phases of SDLC - both in agile and waterfall methodologies
- Practiced scrum methodologies, test-driven, data-driven, behavior-driven environments and CI-CD with Jenkins
- Familiarity with multiple software systems and architecture
- Quick learner, adaptive to new environments, self-motivated, team player and focused. Excellent interpersonal, technical and communication skills
PROFESSIONAL EXPERIENCE
HADOOP DEVELOPER
Confidential, Buffalo, New York
Responsibilities:
- Designed, developed and implemented Hadoop clusters, commissioned and decommissioned nodes and integrated components of Hadoop ecosystem
- Moved data using Sqoop jobs and option files to move date to and HDFS to RDBMSs and vice-versa
- Ingested, processsed and stored on structured, semi-structured and unstructured data on-prem and cloud-based ETL solutions
- Leveraged HDFS commands to manage cluster, manage stored data in HDFS and to apply data compression
- Processed structured data on HDFS using HiveQL, external and managed tables, partitioning, bucketing, joins, UDFs, incremental loads and aggregation
- Optimized Spark performance and leveraged Spark optimizer on existing algorithms to improve performance
- Processed and analysed data by developing Spark applications in Scala, Spark Core, SparkSQL and Spark Streaming
- Applied transformations and actions on RDDs, DataFrames, Paired-RDDs and Datasets to perform business logics
- Processed and analyzed various data of file formats such txt, csv, avro, parquet, orc, JSON, xml and sequence file
- Developed Unix shell scripting
- Developed Scala applications for loading streaming data into NoSQL databases, HBase and HDFS
- Developed data processing, manipulation and analytics application using Scala and Python
- Experienced in RDBMS such as My SQL and No SQL database systems such as HBase
- Identified and resolved performance bottlenecks and managed resources for Spark execution environment
- Experienced in building ETL pipeline using Sqoop and Informatica
- Leveraged compressions such GZip,Bzip2, LZO and Snappy to optimize storage space and network badwidth
- Scheduled data workflow jobs using Oozie and CA Automic to run multiple on-demand and recurring jobs
- Experienced in data modelling, database designs, E-R relationships and performance tuning of MySQL Server
- Tracked project and logged issues with Jira in agile environment
- Worked in a dynamic and collaborative environment with onsite and offshore teammates across different timezones
- Created, peer-reviewed and took ownership of various process documents for clients and other stakeholders
Environment: CDP, HDP, HDFS, Yarn, Sqoop, Hive, HBase, Shell Scripting, Ubuntu, Redhat Linux, Spark Core, Spark SQL, Spark Steaming, Scala, Python, CA Automic, Oozie, Jira, agile scrum
HADOOP DEVELOPER
Confidential, Dallas, Texas
Responsibilities:
- Installed, managed and maintained Hadoop clusters, commissioned and decommissioned nodes and integrated components of Hadoop ecosystem
- Experienced in building ETL pipeline using NiFi, Sqoop and Flume to move data from sources to HDFS
- Prepared, reviewd and updataed design documents for onboarding of clients and other stakeholders
- Written complex Hive Script for joins, map-side joins, incremental loads and aggregation
- Worked with a wide range of file formats including txt, csv, avro, parquet, orc, json and sequence file formats
- Experienced in applying Gzip, Lzop and Snappy compression codes in Hive
- Wrtitten HQL scripts to load data to Hive tables, created partitions and buckets to optimize query perfomance
- Written Hive UDFs as required and executed complex HQLs to extract data from Hive tables
- Handled large datasets using partitions and Spark in-memory capabilities, paired-RDDs, broadcasts, efficient joins, transformations and actions
- Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing
- Implemented Spark UDFs using Scala and SparkSQL for faster testing and processing of data relational data
- Managed datasets using Pandas DataFrames and MySQL, queried MySQL database using Python
- Respobible for managing Spark persistence, resources and memory mangement
- Involved in converting HQL queries into Spark transformations using Spark RDD, Scala.
- Involved in scheduling Oozie workflow to run multiple jobs at a desirable time
- Designed tables and assessed database designs, relationships and performance using MySQL Server
- Imported and exported data with Sqoop to and from HDFS to RDBMS including Oracle, MySQL and MS SQL Server
- Experienced in Eclipse IDE with Maven and SBT frameworks
- Used Jira for agile project management and tracking of bugs and issues
- Writtenshellscripts to automate the process by scheduling and calling the scripts from scheduler.
- Closely collaborated with both the onsite and offshore teams
Environment: Hadoop, HDFS, Nifi, Sqoop, Shell Scripting, Ubuntu, Linux Red Hat, Spark, Scala, Python, Resource Manager, Yarn, Hive
HADOOP DEVELOPER
Confidential, New York, NY
Responsibilities:
- Implemented solution for ingesting data from various sources and processing the data-at-Rest utilizing big data technologies such as Hadoop, hive, Spark, Sqoop, Nifi
- Good knowledge in Using Nifi to automate the data movement between different Hadoop System
- Designed and implemented custom NiFi processors that reacted, processed for the data pipeline A
- Event Streaming on different stages on Stream sets Data Collector, running a MapReduce job on event triggers to convert Avro to Parquet
- Built NiFi dataflow to consume data from Kafka, transform data, store in HDFS and performed spark streaming jobs
- Worked on NoSQL databases including HBase and MongoDB, configured MySQL Database to store Hive metadata.
- Utilized Apache Hadoop environment by Cloudera Hadoop, Hortonworks.
- Configured Zookeeper, Cassandra and Flume to transfer data to Hadoop cluster as per desired schedules.
- Automated all the jobs for pulling data from FTP server and load data into Hive tables, using Oozie workflows.
- Developed complex queries, joins, aggregations and UDFs using Hive and Impala.
- Involved in converting Hive/Sql quries into spark transformation using Spark RDDs.
- Developed MapReduce jobs on Yarn and Hadoop clusters to produce daily and monthly reports.
- Developed a workflow using Oozie to automate the tasks of loading the data into HDFS from analyzing the data.
- Created Phoenix tables, mapped to HBase tables and implemented SQL queries to retrieve data.
- Loaded data from csv files to Spark, created RDDs and DataFrames and queried data using SparkSQL.
- Created external tables in Hive, Loaded JSON format log files and ran queries using HiveQL.
- Involved in designing Avro schemas for serialization and Converting JSON data to Avro format.
- Designed HBase row key and modelling of data to insert into HBase using concepts of lookup and staging tables.
- Used Git for version control of source codes and scripts.
- Created and maintained technical documentation on clusters, tools, technologies and business process mapping
- Optimized Spark performance by improving existing algorithms on SparkSQL, DataFrames, Pair RDDs and YARN
- Leveraged Spark parallelization, partitioning, caching (in-memory, disk, SerDe), Serialization etc. using Scala
- Improved data roadmap and provided solutions to meet new business and technological needs by evaluating existing data platforms, technological stacks and colloborative application of technical expertise.
- Utilized Microsoft data bricks to process spark jobs and blob storage services to process data.
- Worked on data fabrics to process data silos of a big data system.
Environment: Hadoop, HDFS, Kafka, MapReduce, Spark, Hive, Avro, Parquet, Scala, Java, HBase, Cassandra, Hortonworks, ZooKeeper, sqoop, NiFi, Agile scrum
SDET
Confidential, New York, NY
Responsibilities:
- Analyzed technical and functional requirements documents and design and developed QA test plan, test cases and test scenario by maintaining E2E process flow
- Developed testing script for internal brokerage application that is utilized by branch and financial market representatives to recommend and manage customer portfolios; including international and capital markets.
- Designed and Developed Smoke and Regression automation script and Automation of functional testing framework for all modules using Selenium and WebDriver.
- Created data-driven test scripts for adding customers, checking online accounts, user interfaces and reports
- Cross-verified trade entries between mainframe system, web applications and downstream systems
- Extensively used Selenium WebDriver API (XPath and CSS locators) to test the web application.
- Configured Selenium WebDriver, TestNG, Maven tool, Cucumber, and BDD Framework and created Selenium automation scripts in java using TestNG.
- Performed data-driven testing by developing Java based library to read test data from Excel & Properties files.
- Performed queries on DB2 to test and validate the trade entry date from mainframe to backend system.
- Developed data driven framework with Java, Selenium WebDriver and Apache POI to automate a trading system.
- Developed internal application using Angular.js and Node.js connecting to Oracle on the backend.
- Expertise in debugging issues occurred in front end part of web-based application, which is developed using HTML5, CSS3, Angular JS, Node.JS and Java.
- Developed smoke automation test suite for regression test suite.
- Applied various testing technique in test cases to cover all business scenario for quality coverage.
- Interacted with development team to understand design flow, code review, discuss unit test plan.
- Executed tests in System & integration Regression testing In Testing environment.
- Attended triage meeting to discuss on root causes analysis, tracked defect in HP ALM (Quality Center), managed defect by following up open items in Rally, and retest defects with regression testing.
- Provided QA and UAT sign off after closely reviewing all test cases in Quality Center along with receiving the Policy sign off the project
Environment: HP ALM, Selenium WebDriver, JUnit, Cucumber, Angular JS, Node JS, Jenkins, GitHub, Windows, UNIX, Agile, MS SQL, IBM DB2, Putty, WinSCP, FTP Server, Notepad++, C#, DB Visualizer
