Hadoop Developer Resume
Atlanta, GA
SUMMARY:
- Around 6 years of IT experience, diversified exposure in Software Process Engineering, developing, building enterprise applications using Java, SQL and Big Data Technologies.
- Good understanding and implementation knowledge in various domains like Financial sector and Telecom.
- Good understanding and working experience on Hadoop Distributions like Cloudera and Hortonworks.
- Strong working experience with ingestion, storage, querying, processing and analysis of Big data.
- In depth knowledge on Big Data Technologies like HDFS, Map Reduce, Spark, Kafka, Hive, HBase, SQOOP, Scala, Oozie, and Zookeeper.
- Knowledge on working with complex MapReduce programs using Apache Hadoop with different file formats Text, Sequence, Xml, parquet and Avro.
- Worked with different types of datasets like structured, unstructured and semi - structured to analyze data for competitive strategies as per business requirements.
- Worked on developing SPARK Applications using Scala, Spark Core, Spark SQL and Spark Streaming API’s for testing and Fast processing of Data.
- Used Spark RDD’s to convert MapReduce programs into Spark Transformations.
- Working Knowledge on publishing data from external data sources into Kafka topics, and consuming data from Kafka through Spark.
- Involved in ingesting data from servers and load it to the HDFS using Kafka.
- Worked on importing and exporting data using Apache Sqoop from HDFS/Hive/HBase to Relational Database Systems (RDBMS) and vice versa.
- Involved in Creating Hive Tables, Partitioning, Bucketing, loading with data and Writing complex HIVE HQL quires to extract data as per requirements.
- Written complex HIVE quires for processing and analyzing large volumes set of data and use the resulting knowledge for various business benefits, opportunities for new revenues and gain a competitive advantage.
- Implemented automatic workflows and job scheduling using Oozie,
- Working knowledge on NOSQL Databases .
- Good Working knowledge on JAVA, used Eclipse IDE and Maven for developing and debugging JAVA Applications.
- Familiar with AWS Components like EC2, S3, EMR.
- Experience in using version control tools like GITHUB to share the code snippet among the team members
- Experience on Tableau for making reports on data.
- Experience in working in Agile project management.
- Involved in daily SCRUM meetings to discuss the development/progress and was active in making scrum meetings more productive.
- Strong motivational skills, familiarity with various technologies, ability to learn quickly, dealing with people, commitment to work and believes in hard-working.
TECHNICAL SKILLS:
Big Data Technologies\Processing\: HDFS, MapReduce, Spark, Kafka, Hive, \Spark, Spark Streaming, Spark SQL, Hive, \ Sqoop, Zookeeper, Oozie, HBase.\MapReduce, EMR, Ec2.\
Big Data Platforms\ETL\: Cloudera, Hortonworks.\Kafka, Sqoop\
Databases\Languages\: MySQL, SQL Server, HBase, Hive.\Java, Scala, SQL, \
Methodologies\BI Tools\: Agile\Tableau \
Bug Tracking Tools\: Jira\
PROFESSIONAL EXPERIENCE:
Confidential, Atlanta GA
Hadoop Developer
Responsibilities:
- Responsible for ingesting user behavioral data and customer profile data to Analytics Data store on daily basis.
- Developed Spark programs to analyze and transform the data.
- Run Sqoop jobs to migrate data from MySQL Database to Hadoop File System.
- Build data pipeline using Kafka and developed Kafka producers to stream data from external rest API’s to Kafka Topics.
- Developed Kafka Consumer API’s in Scala for consuming data from Kafka topics and loaded processed streams to HIVE for future analysis.
- Developed Spark applications using Scala for performing data cleansing, event enrichment, data aggregation, de-normalization and data preparation needed for reporting teams to consume.
- Experienced in handling large datasets using Spark in Memory capabilities, using broadcasts variables in Spark, effective & efficient Joins, transformations and other capabilities.
- Used Scala based Spark RDD’s and Spark- SQL for faster testing and processing of ETL workloads and loaded the results into HIVE.
- Involved in creating HIVE tables, imported processed data into tables and done analytics as per business developments.
- For performance optimization implemented dynamic partitions, bucketing and compression techniques in Hive External Tables for fast query processing.
- Worked on troubleshooting and Performance optimization on Spark Application to improve the over-all processing time for the pipelines.
- Used MySQL for storing customer profile data, moved data to HDFS when needed and run Spark jobs for analyzing customer behavior.
- Generated daily final report’s on processed data using Tableau.
- Used JIRA tracking tool for assigning and defect management.
- Worked in Agile development environment and actively involved in daily Scrum and other design related meetings.
Environment: Hortonworks (HDP), Hadoop, HDFS, Hive, Spark, Scala, Kafka, Sqoop, MySQL
Confidential, New York, NY
HADOOP DEVELOPER
Responsibilities:
- Responsible for building scalable distributed data solutions usingHadoop.
- Handled importing of data from various data sources, performed transformations using Hive and loaded data into HDFS for aggregations.
- Extracted transactional data from MySQL Database and loaded the data into HDFS file system using Sqoop.
- Developed simple to complex Map/Reduce jobs using Java, and scripts using Hive.
- Involved in creating Hive Tables and load the processed data into tables.
- Worked on various performance optimizations like using distributed cache for small datasets, Partition,Bucketing in Hive and Map Side joins.
- Analyzed the large data sets by performing Hive queries (HiveQL) and stored the results in Hadoop Data File System.
- Implemented Spark using Scala and SparkSQL for faster testing and processing of data.
- Scheduled and executed workflows in Oozie to run Hive.
- Experienced on loading and transforming of large sets of structured and semi structured data.
- Designed and developed Dashboards for Analytical purposes using Tableau.
Environment: Cloudera(CDH), Hadoop, HDFS, Map Reduce, Spark, Scala, HIVE, MYSQL, Sqoop, Oozie
Confidential
Java Developer
Responsibilities:
- Involved in all the phase of project development - requirements gathering, analysis, design, development, coding and testing.
- Used Maven script for building and deploying the application.
- Strong experience in SQL and knowledge of MySQL Server database
- Wrote SQL quires to fetch the statistics from the database and displayed the statistics in the application UI.
- Experience in ETL software development.
- Involved in code review and documentation review of technical artifacts.
- Used Tortoise SVN for source code versioning and code repository.
- Used Jenkins for Continuous Integration Builds and deployments(CI/CD)
- Used JIRA to track bugs.
Environment: Java, MySQL, Maven 3.2, JIRA, Jenkins, Agile, Eclipse IDE
Confidential
Java Application Developer.
Responsibilities:
- Involved in designing and programming the OBS systems.
- Designed Development Cycle by creating Flow Diagram, Entity Relationship Diagram, Data Flow Diagram and Database Design.
- Created tables, stored procedures in SQL for data manipulation and retrieval, Database Modification using SQL, Stored procedures in MySQL.
- Developed Java APIs, which communicates with the Java Beans.
- Used Controllers for generating customized reports, transaction, logins and reporting modules.
- Developed POJO classes and wrote SQLQuires for data extraction.
- For Data Presentation, Report Generation and customer feedback used XML and XSL Documents.
- Used Java Beans to automate the generation of dynamic Reports and customer Transactions.
- Used Log4J for logging throughout the application.
- Used JIRA to track bugs.
Environment: Java, JavaBeans, JDBC, MySQL, XML, XSL, SVN, Log4J, Eclipse IDE
