Hadoop Developer Resume
AtlantA
SUMMARY
- A software professional with over 5+ years of experiencing developing applicationsin the field of information technology with extensive knowledge on Hadoop Eco System.
- Hands on experience in Hadoop eco - system technologiesApache Spark, Hive, Pig, Sqoop,Oozie, Zookeeper, Kafka, Storm, Apache HBase and MapReduce/YARN, HDFS.
- Hands-on experience on full life cycle implementation using HDP (HORTONWORKS) distributions and good knowledge on CDH(Cloudera).
- Strong Knowledge on Architecture of Distributed systems and Parallel processing, In-depth understanding of MapReduce programing paradigm and Spark execution framework.
- Expertise in writing Scala APIs for data analysis and data manipulation by creating data frames in Spark.
- Strong knowledge on HiveQL and experience in writing queries to analyze data stored across multiple file formats such as text, Sequence files, Avro, ORC, and Parquet.
- Experience in writing UDFs for Hive, Pig, custom MapReduce programs.
- Worked on creating Sqoop jobs to import data from relational databases into HDFS and exporting analyzed data into relational databases for report generation and visualization.
- Good knowledge using Spark Framework for batch and real-time data processing using Spark Streaming.
- Experienced in tuning long running Spark applications and implementing features like graceful shutdown, fault tolerance and fail over.
- Experience in migrating unstructured datainto the HDFS using Kafka.
- Hands on experience in non-relational databases like HBase and MongoDB.
- Experienced in data ETL from SQL/NoSQL databases and worked on integrating data from various file formats into a combined analytical data model for reporting.
- Experiencewith versioncontrolsystemslikeGitHub and SVN.
- Good knowledge SDLC, Agile and WaterfallMethodologies.
- Involved in dailySCRUMmeetingstodiscussthedevelopment/progressand wasactivein making scrum meetings moreproductive.
- Experience in analysis, requirement specification and design and solution description.
- Team player with good interpersonal skills, strong understanding of fundamental business processes and Communication skills.
TECHNICAL SKILLS
Operating Systems: Windows Server 2000/2003/2008, Windows 10/8/7/XP/Vista, MAC OS, Linux (Red Hat, Ubuntu),cent OS 6.0
Hadoop Framework: HDFS, Map Reduce and YARN, Spark, HIVE, PIG, Oozie, Zookeeper.
Languages: JAVA, Scala, HiveQL, Pig-Latin, Linux Commands and SQL.
Databases: HBase, MongoDB, Oracle, MySQL, MS SQL Server.
Data Ingestion / ETL Tools: Sqoop, Storm, Flume, Kafka.
IDE: Eclipse and IntelliJ
PROFESSIONAL EXPERIENCE
Confidential, Atlanta
Hadoop Developer
Responsibilities:
- Worked on HortonworksHadoop distribution systems.
- Performed various data validations, data cleansing and data aggregation using series of Spark Applications written using Scala.
- Responsible for designing and implementing the data pipeline using Big Data tools including Hive, Oozie, Spark, Drill, Sqoop.
- Written scripts in HiveQL and pig-Latin to analyze the data and build test cases to test and detect early defects in the imported data.
- Worked with different file formats like ORC, Parquet,Avro and compression techniques like gzip and Snappy.
- Written queries usingHiveQL to analyze the data and Provided reports and data metrics to different LOBs leveraging Hive.
- Worked with application developers and project stakeholders to perform testing, troubleshooting of the applications to provide the best possible output before boarding to Hadoop.
- Worked on HBase for manipulation of structured, semi-structured and unstructured data.
- Developed Oozie Workflows to string all the actions to develop end to end data analytic pipelines.
- Developed UDFs in JAVA to performMapReduce jobs.
- Worked with multiple individuals from different levels of SDLC to coordinate, prioritize, schedule and track project goals.
Environment: Hortonworks Hadoop Platform, Spark, Hive, Pig, HDFS, Map Reduce, Sqoop, Kafka, HBase, Oozie, Zookeeper, Maven, MySQL, Eclipse, Java and Agile methodologies.
Confidential, Miami, FL.
Hadoop Developer
Responsibilities:
- Responsible for loading and managing of data in the Hadoop cluster.
- Created Scala scripts using Spark and SparkSQL to manipulate and analyze the data.
- Wrote scripts using HiveQL and Pig-Latin to analyze the web log data of the customers.
- Worked on Pig Latin scripts to identify and extract the data from the web server output files to load into HDFS.
- Used Kafka and Stromforprocessingapplication log data in realtime.
- Setting up theStrom Streaming and Kafka Cluster.
- Prioritized Hive queries, Pig scripts, Sqoop jobs and Map-Reduce Programs using Oozie workflows and sub-workflows.
- Configuring Zookeeper to provide various cluster coordination services.
- Coordinated with QA team during testing phase.
Environment: Hortonworks Hadoop platform, Linux, Spark, Hive, Pig, MapReduce/YARN,HDSF, Sqoop, Flume, Oozie, HBase, Maven, JAVA, Eclipse, Putty, Shell Scripting and MySQL.
Confidential
Hadoop Developer
Responsibilities:
- Worked onanalyzing Hadoop cluster and different Big Data analytictools including Pig, Hive, HBase andSqoop.
- Installed Hadoop, MapReduce, HDFS, and developed multiple MapReduce jobs in Pig and Hive for data cleaning andpre-processing.
- Coordinated with business customers to gather business requirements, also interacted with other technical peers to derive technical requirements and delivered the BRD and TDD documents.
- Involved in upgrading Hadoop Cluster from HDP 1.3 to HDP 2.0.
- Created HBase tables to store various data formats of PII data coming from different portfolios.
- Involved in Testing and coordination with business in User Testing.
- Populate HDFS with huge amounts of data using Apache Kafka.
- Used Hive to analyze the partitioned and bucketed data to compute various metrics for reporting.
- Worked on POCs using Kafka and Spark for real time data processing.
- Developed Hadoop streaming Map/Reduce work using Java.
Environment: Hadoop, Hive, MapReduce, Pig, MongoDB, Oozie, Sqoop, Kafka, Spark, HBase, HDFS, Javan and Zookeeper.
Confidential
Jr. SQL Developer
Responsibilities:
- Participating client Business call to understand new requirements.
- Providing transition to team to understand the flow of data to downstream.
- Performed Database Testing to check data validity and integrity on SQL Server.
- Making PL/SQL code change as per new Business Requirement.
- Worked on creating anonymous PL/SQL code.
- Working with team, providing them the process document to understand applications.
- Writing complex SQL queries to check data in database.
- Implementing anonymous Pl/SQL block to check data.
Environment: TOAD, SQL Developer, MSSQL.
