We provide IT Staff Augmentation Services!

Hadoop /spark Bigdata Developer Resume

3.00/5 (Submit Your Rating)

Richfield, MN

SUMMARY

  • Overall 10+ years of experience in design and deployment of Data Management and Data Warehousing Projects in various roles as a Data Modeler and Data Analyston Big data technologies.
  • Possesses 3+ years of rich Hadoop experience in design and development of Big Data applications, which involves Apache Hadoop Map/Reduce, HDFS, Hive, HBase, Pig, Oozie, Sqoop, Flume and Spark.
  • Expertise in developing solutions around NOSQL databases like MongoDB and Cassandra.
  • Experience with all flavor of Hadoop distributions, including Cloudera, Horton works.
  • Excellent understanding of Hadoop architecture Map Reduce MRv1 and Map Reduce MRv2 (YARN).
  • Developed multiple Map Reduce programs to process large volumes of semi/unstructured data files using different Map Reduce design patterns.
  • Strong experience in writing Map Reduce jobs in Java and Pig.
  • Experience with various performance optimizations like using distributed cache for small datasets, partition, bucketing in Hive and Map Side joins when writing Map Reduce jobs.
  • Excellent understanding of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
  • Worked extensively over semi - structured data (fixed length & delimited files), for data sanitation, report generation and standardization.
  • Excellent hands on experience in analyzing data using Pig Latin, HQL, HBase and Map Reduce programs in Java.
  • Developed UDF's in Java as and when necessary to use with PIG and HIVE queries.
  • Have dealt with Zookeeper an Oozie Operational Services for coordinating the cluster and scheduling workflows.
  • Strong Knowledge of Hadoop and Hive and Hive's analytical functions.
  • Loaded the dataset into Hive for ETL Operation.
  • Proficient using of big data ingestion tools like Flume and Sqoop.
  • Experience in importing and exporting data between HDFS and Relational Database Management systems using Sqoop.
  • Experience in handling continuous streaming data using Flume and memory channels.
  • Good experience in benchmarking Hadoop cluster.
  • Good knowledge on data analysis with SAS.
  • Good knowledge on executing Spark SQL queries against data in Hive.
  • Experienced in monitoring Hadoop cluster using Cloudera Manager and Web UI.
  • Experience in implementing setting up standards and processes for Hadoop based application design and implementation.
  • Extensive Experience working on web technologies like HTML, CSS, XML, JSON, JQuery.
  • Hands-on experience with AWS (Amazon Web Services), using Elastic Map Reduce (EMR), creating and Storing data in S3 buckets and creating Elastic Load Balancers (ELB) for Hadoop front end Web UI’s.
  • Extensive knowledge on creating Hadoop cluster on multiple EC2 instances in AWS and configuring them through ambari and using IAM (Identity and Access Management) for creating groups, users.
  • Extensive experience in documenting requirements, functional specifications and technical specifications.
  • Extensive experience with SQL, PL/SQL and database concepts.
  • Experience working on Version control tools like SVN and GIT revision control systems such as GitHub and JIRA to track issues and crucible for code reviews.
  • Strong Database background with Oracle, PL/SQL, Stored Procedures, trigger, SQL Server, MySQL.
  • Strong Problem Solving and Analytical skills and abilities to make Balanced & Independent Decisions.
  • Good Team Player, Strong Interpersonal, Organizational and Communication skills combined with Self-Motivation, Initiative and Project Management Attributes.
  • Holds strong ability to handle multiple priorities and work load and also has ability to understand and adapt to new technologies and environments faster.

TECHNICAL SKILLS

Hadoop Core Services: HDFS, Map Reduce, Spark, YARN

Hadoop Distribution: Horton works, Cloudera, Apache

NO SQL Databases: MongoDB, Cassandra

Hadoop Data Services: Hive, Pig, Sqoop, Flume, Sqoop

Hadoop Operational Services: Zookeeper, Oozie

Monitoring Tools: Ambari, Cloudera Manager

Cloud Computing Tools: Amazon AWS

Languages: C, Java, Python, SQL, PL/SQL, Pig, HiveQL, Unix Shell Scripting

Databases: Oracle, MySQL, MongoDB

Operating Systems: UNIX, Windows, LINUX

Build Tools: Jenkins, Maven, ANT

Development Tools: Microsoft SQL Studio, Toad, Eclipse

Development Methodologies: Agile/Scrum, Waterfall

PROFESSIONAL EXPERIENCE

Hadoop /Spark Bigdata Developer

Confidential, Richfield, MN

Responsibilities:

  • Worked on Sqoop jobs for ingesting data from Oracle and MySQL
  • Created hive external tables for querying the data
  • Used Spark Data frame APIs to ingest S3 data
  • Wrote scripts to load data from Red shift
  • Processed complex/nested json and csv data using Data frame API
  • Applied Transformation rules on the top of Data Frames
  • Scheduled Spark jobs using Oozie
  • Processed Hive, csv, json and oracle data
  • Validated the source and final output data.
  • Tested the data using Dataset API
  • Partitioned (dynamic as well as static partition) and Bucketed tables to improve query performance
  • Improved HQL performance by analyzing the plan using explain plan and applying various optimization techniques like Map side join, join optimization, tuning container (CPU/Core, Memory etc.)
  • Based on new spark versions, applying different optimization transformation rules
  • Debugged the script to minimize shuffling of data
  • Analyzed and created reports using Tableau
  • Created dashboards in Tableau

Environment: Hadoop, Spark/Scala, MapReduce, HDFS, HBase, Hive, Pig, Java, SQL, Sqoop, Flume, Oozie, UNIX, Maven, Eclipse

Big Data Hadoop Developer

Confidential

Responsibilities:

  • Installed/Configured/Maintained ApacheHadoopclusters for application development andHadoop tools likeHive, Pig, and Sqoop
  • Configured, Designed implemented and monitored Kafka cluster and connectors
  • Used Sqoop to import data into HDFS and Hive from multiple data systems
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.Handled importing of data from various data sources, performed transformations using Hive, MapReduce
  • Helped with the sizing and performance tuning of the Cassandra cluster
  • Involved in converting Cassandra/Hive/SQL queries into Spark transformations using Spark RDD's
  • Developed multiple POCs using Spark and deployed on the Yarn cluster
  • Involved in the process of Cassandra data modeling and building efficient data structures
  • Extracted the data from Teradata into HDFS using Sqoop
  • Analyzed the data by performing Hive queries and running Pig scripts to know user behavior like shopping
  • Configured Oozie workflow to run multiple Hive and Pig jobs which run independently with time and data availability
  • Optimized MapReduce code, pig scripts and performance tuning and analysis
  • Implemented advanced procedures like text analytics and processing, using the in-memory computing capabilities of Spark
  • Exported the aggregated data into Oracle using Sqoop for reporting on the Tableau dashboard
  • Involvement in design, development and testing phases of Software Development Life Cycle
  • Performedinstallation, updates, patches and version upgrades when required for Hadoop

Environment: Hadoop, Map Reduce, HDFS, HBase, Hive, Pig, Java, SQL, Sqoop, Flume, Oozie, UNIX, Java, Maven, Eclipse

Data Integration Developer/Analyst

Confidential

Responsibilities:

  • Involved in gathering specifications from Business User and designing of the process and set up time lines for entire process
  • Extracting updated data Portals periodicallywith SAS/SQL
  • Coordinating with different teams to make sure data is available on time
  • Cleansing and validating data
  • Creating data sets for analysis and report.
  • Responsible to maintain previous month and previous financial year
  • Responsible to validate data reports with historical data.
  • Responsible for doing UAT before release to business users
  • Performed documentation of the SAS Code for the better understanding of the program
  • Involved in automation of reports in various formats like pdf, html and excel reports.
  • Extensively used procedures like PROC SQL, PROC PRINT, and PROC SORTetc.
  • Coded SAS programs with the use ofSAS jobs.
  • Analyzed the data using SAS/STAT proceduresPROC FREQ,PROC MEANS,PROC

Environment: SAS, Oracle, SQL, Linux

Database Analyst/Developer

Confidential

Responsibilities:

  • Work with Customer Analytics and key business stakeholders to prioritize the order in which disparate customer data sources are integrated into the customer data mart
  • Provided analysis for senior management regarding ROI of promotions leading to shift marketing budget into the most effective channels
  • Utilized advanced analytical methods in SAS and Microsoft Excel, including marketing mix models, to test the effectiveness of Hershey's promotional activities
  • Employed forecasting models to understand underlying sales trends and expected future performance
  • Monitor database performance, implement required changes

Environment: SAS, MS SQL Server, Excel, Windows

Java Developer

Confidential

Responsibilities:

  • Used Object Oriented Programming and design.(OOP&OOD)
  • Wrote stored procedures, complex queries using PL/SQL to extract data from the database, delete data and reload data on Oracle9i DB using the Toad tool.
  • Developed both front-end and back-end of the product usingJava, J2EE, Ajax, JQuery, spring and Hibernate, and other technologies.
  • Developed user interfaces using JSPs, HTML, CSS,JavaScript, jQuery, JSPCustomTags.
  • Used Spring Core Annotations for Dependency Injection.

Environment: Java, J2EE, JSP, spring, Hibernate, Agile, Tomcat, Web Services, MySQL, Eclipse 3.5, SVN, Maven, JUnits, Hudson, JMS

We'd love your feedback!