Senior Hadoop Developer Resume
Cincinnati, OH
SUMMARY:
- 8 + years of IT industry experience with 5 years of experience in dealing with Apache Hadoop components like HDFS, MapReduce, Spark, Hive, Pig, Sqoop, Oozie, Zookeeper, HBase, Cassandra, MongoDB and Amazon Web Services.
- 3+ years of experience in the Application Development and Maintenance of SDLC projects using Java technologies.
- Good experience working with Hortonworks Distribution, Cloudera Distribution and MapR Distribution
- Very good understanding/knowledge of Hadoop Architecture and various components such as HDFS, JobTracker, TaskTracker, NameNode, DataNode, Secondary Namenode, and MapReduce concepts.
- Developed applications for Distributed Environment using Hadoop, Mapreduce and Python.
- Experience in data extraction and transformation using MapReduce jobs.
- Proficient in working with Hadoop, HDFS, writing PIG scripts and Sqoop scripts.
- Performed data analysis using Hive and Pig.
- Expert in creating Pig and Hive UDFs using Java in order to analyze the data efficiently.
- Experience in importing and exporting data using Sqoop from Relational Database Systems to HDFS and vice - versa.
- Strong understanding of NoSQL databases like HBase, MongoDB & Cassandra.
- Strong understanding of Spark real time streaming and SparkSQL and experience in loading data from external data sources like MySQL and Cassandra for Spark applications.
- Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
- Well versed with job workflow scheduling and monitoring tools like Oozie
- Developed MapReduce jobs to automate transfer of data from HBase.
- Practical knowledge on implementing Kafka with third-party systems, such as Spark and Hadoop.
- Loaded streaming log data from various webservers into HDFS using Flume.
- Experience in using Sqoop, Oozie and Cloudera Manager.
- Hands on experience in application development using RDBMS, and Linux shell scripting.
- Have experience with working on Amazon EMR and EC2 Spot instances
- Experience in integrating Hadoop with Ganglia and have good understanding of Hadoop metrics and visualization using Ganglia.
- Support development, testing, and operations teams during new system deployments.
- Solid understanding of relational database concepts.
- Extensively worked with Unified Modeling Tools (UML) in designing Use Cases, Activity flow diagram, Class diagrams, Sequence and Object Diagrams using Rational Rose, MS-Visio.
- Very good experience in complete project life cycle (design, development, testing and implementation) of Client Server and Web applications.
- Good Knowledge on Hadoop Cluster architecture and monitoring the cluster.
- Hands on experience in Tableau to generate Hadoop data report.
- Good team player and can work efficiently in multiple team environments and multiple products. Easily adaptable to the new systems and environments.
- Possess excellent communication and analytical skills along with a can-do attitude.
WORK EXPERIENCE:
Senior Hadoop Developer
Confidential, Cincinnati, OH
Responsibilities:
- Strong understanding and practical experience in developing Spark applications with Scala.
- Developed Spark scripts by using Spark shell commands as per the requirement.
- Developed Scala scripts, UDF's using both Data frames/SQL and RDD in Spark for Data Aggregation.
- Exploring with Spark for improving the performance and optimization of the existing algorithms in Hadoop using Spark context, Spark-SQL, Data Frame and pair RDD's
- Experience in developing SparkSQL applications both using SQL and DSL
- Extensively worked with parquet file format and gained practical knowledge in writing spark and hive applications to meet the parquet requirements.
- Experience in using various compression techniques along with Parquet file format.
- Experience in managing extensive retail datasets from Kroger and gained good experience in creating the test datasets for development purpose
- Experience in building dimensional and fact tables using Spark Scala applications
- Practical knowledge on writing applications in Scala to interact with the Hive through the Spark application.
- Extensively used Hive partitioned tables, map join, bucketing and gained good understanding of dynamic partitioning.
- Performed POC on writing the spark applications in Scala, Python and R programming language
- Good hands on experience with Hive to perform data queries and analysis as a part of the QA
- Practical experience in using Pig to perform the QA by calculating the statistics of the final output.
- Experience in designing both time driven and data driven automated workflows using Oozie
- Experience in writing Sqoop scripts to import data from exadata to HDFS
- Good exposure to MongoDB, it's functionality and use-cases
- Gained good exposure to Hue interface for monitoring the job status, managing the HDFS files, tracking the scheduled jobs and managing the Oozie workflows
- Performed optimizations and performance tuning in Spark and Hive
- Developed Unix script to automate data load into HDFS
- Strong knowledge on HDFS commands to manage the files and also gained good understanding in managing the file system through the Spark Scala applications.
- Extensive usage of alias for Oozie and HDFS commands
- Experienced in managing and reviewing Hadoop log files.
- Experience in log controlling for Spark applications and extensive use of log4j to log the respective phases of the application accordingly
- Good knowledge on GIT commands, version tagging and pull requests
- Performed unit testing and also integration testing after the development and participated in code reviews.
- Experience in writing the Junit test cases for testing the Spark and SparkSQL applications
- Practical experience with developing applications in IntelliJ and Maven
- Good exposure to Agile environment. Participated in daily standups, Big Room Planning, Sprint meetings and Team Retrospectives
- Interact with business analysts to understand the business requirements and translate them to technical requirements
- Collaborate with various technical experts, architects and developers for design and implementation of technical requirements
- Documented business requirements, technical specifications, and process flows.
Environment: Hadoop 2.6.0-cdh5.7.0, Java 1.8.0 92, Spark 1.6.0, SparkSQL, R programming, Python, Scala 2.10.5, MongoDB, Apache Pig 0.12.0, Apache Hive 1.1.0, HDFS, Sqoop, Oozie, Maven, IntelliJ, GIT, UNIX Shell scripting, Oracle 11g/10g, Log4j, Linux, Agile development
Hadoop Developer
Confidential, Memphis, TN
Responsibilities:
- Extracted and updated the data into HDFS using Sqoop import and export command line utility interface.
- Responsible for developing data pipeline using Flume, Sqoop, and Pig to extract the data from weblogs and store in HDFS.
- Involved in developing Hive UDFs for the needed functionality.
- Involved in creating Hive tables, loading with data and writing Hive queries.
- Managed works including indexing data, tuning relevance, developing custom tokenizes and filters, adding functionality includes playlist, custom sorting and regionalization with Solr search engine.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Used pig to do transformations, event joins, filter boot traffic and some pre-aggregations before storing the data onto HDFS.
- Implemented advanced procedures like text analytics and processing using the in-memory computing capabilities like spark.
- Enhanced and optimized product Spark code to aggregate, group and run data mining tasks using the Spark framework.
- Extending Hive functionality by writing custom UDFs.
- Experience in managing and reviewing Hadoop log files
- Developed data pipeline using Flume, Sqoop, pig and java MapReduce to ingest customer behavioral data and financial histories into HDFS for analysis.
- Involved in emitting processed data from Hadoop to relational databases and external file systems using Sqoop.
- Orchestrated hundreds of Sqoop scripts, pig scripts, Hive queries using Oozie workflows and sub-workflows.
- Loaded cache data into HBase using Sqoop.
- Experience in custom talend jobs to ingest, enrich and distribute data in MapR, Cloudera Hadoop ecosystem.
- Created lots of external tables on Hive pointed to HBase tables.
- Analyzed HBase data in Hive by creating external partitioned and bucketed tables.
- Worked with cache data stored in Cassandra.
- Injected the data from External and Internal Flow Organizations.
- Used the external tables in Impala for data analysis.
- Supported MapReduce Programs those are running on the cluster.
- Participated in apache Spark POCS for analyzing the sales data based on several business factors
- Participated in daily scrum meetings and iterative development.
Environment: Hadoop, MapReduce, Hdfs, Pig, Hive, HBase, Impala, Sqoop, Flume, Oozie, Apache Spark, Java, Linux, SQL Server, Zookeeper, Autosys, Tableau, Cassandra.
Hadoop Developer
Confidential, Franklin Lakes, NJ
Responsibilities:
- Worked with systems engineering team to plan and deploy new Hadoop environments and expand existing Hadoop clusters with agile methodology.
- Monitored multiple Hadoop clusters environments using Ganglia, monitored workload, job performance and capacity planning using Cloudera Manager.
- Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
- Experienced with through hands-on experience in all Hadoop, Java, SQL and Python.
- Used Flume to collect, aggregate, and store the web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
- Participated in functional reviews, test specifications and documentation review
- Performed MapReduce programs on log data to transform into structured way to find user location, age group, spending time.
- Analyzed the web log data using the HiveQL to extract number of unique visitors per day, page views, visit duration, most purchased product on website.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports by Business Intelligence tools.
- Oozie with the rest of the Hadoop stack supporting several types of Hadoop jobs out of the box (such as MapReduce, Pig, Hive, and Sqoop) as well as system specific jobs (such as Java programs and shell scripts).
- Proactively monitored systems and services, architecture design and implementation of Hadoop deployment, configuration management, backup, and disaster recovery systems and procedures.
- Experience using Talend for ETL tools and also extensive knowledge on Netezza.
- Installed and configured Flume, Hive, Pig, Sqoop and Oozie on the Hadoop cluster, involved in Analyzing system failures, identifying root causes, and recommended course of actions.
- Documented the systems processes and procedures for future references, responsible to manage data coming from different sources.
Environment: Hadoop, HDFS, Map Reduce, Flume, Pig, Sqoop, Hive, Pig, Sqoop, Oozie, Gangli
Software Engineer
Confidential, Dover, NH
Responsibilities:
- Developed the user interface screens using swing for accepting various system inputs such as contractual terms, monthly data pertaining to production, inventory and transportation.
- Involved in designing database connections using JDBC.
- Involved in design and development of UI using HTML, JavaScript and CSS.
- Involved in creating tables, stored procedures in Sqlfor data manipulation and retrieval using sqlsever2000, database modification using Sql, Pl/Sql, triggers, views in oracle.
- Used dispatch action to group related actions into a single class.
- Build the applications using Ant tool, also used eclipse as the IDE.
- Developed the business components used for the calculation module.
- Involved in the logical and physical database design and implemented it by creating suitable tables, views and triggers.
- Applied J2EE design patterns like business delegate, DAO and singleton.
- Created the related procedures and functions used by JDBC calls in the above requirements.
- Actively involved in testing, debugging and deployment of the application on WebLogic application server.
- Developed test cases and performed unit testing using JUnit.
- Involved in fixing bugs and minor enhancements for the front-end modules.
Environment: Java, HTML, Java script, CSS, Oracle, JDBC, ANT tool, SQL, Swing and Eclipse.
