Hadoop /spark Bigdata Developer Resume
Richfield, MN
SUMMARY
- Overall 10+ years of experience in design and deployment of Data Management and Data Warehousing Projects in various roles as a Data Modeler and Data Analyston Big data technologies.
- Possesses 3+ years of rich Hadoop experience in design and development of Big Data applications, which involves Apache Hadoop Map/Reduce, HDFS, Hive, HBase, Pig, Oozie, Sqoop, Flume and Spark.
- Expertise in developing solutions around NOSQL databases like MongoDB and Cassandra.
- Experience with all flavor of Hadoop distributions, including Cloudera, Horton works.
- Excellent understanding of Hadoop architecture Map Reduce MRv1 and Map Reduce MRv2 (YARN).
- Developed multiple Map Reduce programs to process large volumes of semi/unstructured data files using different Map Reduce design patterns.
- Strong experience in writing Map Reduce jobs in Java and Pig.
- Experience with various performance optimizations like using distributed cache for small datasets, partition, bucketing in Hive and Map Side joins when writing Map Reduce jobs.
- Excellent understanding of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
- Worked extensively over semi - structured data (fixed length & delimited files), for data sanitation, report generation and standardization.
- Excellent hands on experience in analyzing data using Pig Latin, HQL, HBase and Map Reduce programs in Java.
- Developed UDF's in Java as and when necessary to use with PIG and HIVE queries.
- Have dealt with Zookeeper an Oozie Operational Services for coordinating the cluster and scheduling workflows.
- Strong Knowledge of Hadoop and Hive and Hive's analytical functions.
- Loaded the dataset into Hive for ETL Operation.
- Proficient using of big data ingestion tools like Flume and Sqoop.
- Experience in importing and exporting data between HDFS and Relational Database Management systems using Sqoop.
- Experience in handling continuous streaming data using Flume and memory channels.
- Good experience in benchmarking Hadoop cluster.
- Good knowledge on data analysis with SAS.
- Good knowledge on executing Spark SQL queries against data in Hive.
- Experienced in monitoring Hadoop cluster using Cloudera Manager and Web UI.
- Experience in implementing setting up standards and processes for Hadoop based application design and implementation.
- Extensive Experience working on web technologies like HTML, CSS, XML, JSON, JQuery.
- Hands-on experience with AWS (Amazon Web Services), using Elastic Map Reduce (EMR), creating and Storing data in S3 buckets and creating Elastic Load Balancers (ELB) for Hadoop front end Web UI’s.
- Extensive knowledge on creating Hadoop cluster on multiple EC2 instances in AWS and configuring them through ambari and using IAM (Identity and Access Management) for creating groups, users.
- Extensive experience in documenting requirements, functional specifications and technical specifications.
- Extensive experience with SQL, PL/SQL and database concepts.
- Experience working on Version control tools like SVN and GIT revision control systems such as GitHub and JIRA to track issues and crucible for code reviews.
- Strong Database background with Oracle, PL/SQL, Stored Procedures, trigger, SQL Server, MySQL.
- Strong Problem Solving and Analytical skills and abilities to make Balanced & Independent Decisions.
- Good Team Player, Strong Interpersonal, Organizational and Communication skills combined with Self-Motivation, Initiative and Project Management Attributes.
- Holds strong ability to handle multiple priorities and work load and also has ability to understand and adapt to new technologies and environments faster.
TECHNICAL SKILLS
Hadoop Core Services: HDFS, Map Reduce, Spark, YARN
Hadoop Distribution: Horton works, Cloudera, Apache
NO SQL Databases: MongoDB, Cassandra
Hadoop Data Services: Hive, Pig, Sqoop, Flume, Sqoop
Hadoop Operational Services: Zookeeper, Oozie
Monitoring Tools: Ambari, Cloudera Manager
Cloud Computing Tools: Amazon AWS
Languages: C, Java, Python, SQL, PL/SQL, Pig, HiveQL, Unix Shell Scripting
Databases: Oracle, MySQL, MongoDB
Operating Systems: UNIX, Windows, LINUX
Build Tools: Jenkins, Maven, ANT
Development Tools: Microsoft SQL Studio, Toad, Eclipse
Development Methodologies: Agile/Scrum, Waterfall
PROFESSIONAL EXPERIENCE
Hadoop /Spark Bigdata Developer
Confidential, Richfield, MN
Responsibilities:
- Worked on Sqoop jobs for ingesting data from Oracle and MySQL
- Created hive external tables for querying the data
- Used Spark Data frame APIs to ingest S3 data
- Wrote scripts to load data from Red shift
- Processed complex/nested json and csv data using Data frame API
- Applied Transformation rules on the top of Data Frames
- Scheduled Spark jobs using Oozie
- Processed Hive, csv, json and oracle data
- Validated the source and final output data.
- Tested the data using Dataset API
- Partitioned (dynamic as well as static partition) and Bucketed tables to improve query performance
- Improved HQL performance by analyzing the plan using explain plan and applying various optimization techniques like Map side join, join optimization, tuning container (CPU/Core, Memory etc.)
- Based on new spark versions, applying different optimization transformation rules
- Debugged the script to minimize shuffling of data
- Analyzed and created reports using Tableau
- Created dashboards in Tableau
Environment: Hadoop, Spark/Scala, MapReduce, HDFS, HBase, Hive, Pig, Java, SQL, Sqoop, Flume, Oozie, UNIX, Maven, Eclipse
Big Data Hadoop Developer
Confidential
Responsibilities:
- Installed/Configured/Maintained ApacheHadoopclusters for application development andHadoop tools likeHive, Pig, and Sqoop
- Configured, Designed implemented and monitored Kafka cluster and connectors
- Used Sqoop to import data into HDFS and Hive from multiple data systems
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.Handled importing of data from various data sources, performed transformations using Hive, MapReduce
- Helped with the sizing and performance tuning of the Cassandra cluster
- Involved in converting Cassandra/Hive/SQL queries into Spark transformations using Spark RDD's
- Developed multiple POCs using Spark and deployed on the Yarn cluster
- Involved in the process of Cassandra data modeling and building efficient data structures
- Extracted the data from Teradata into HDFS using Sqoop
- Analyzed the data by performing Hive queries and running Pig scripts to know user behavior like shopping
- Configured Oozie workflow to run multiple Hive and Pig jobs which run independently with time and data availability
- Optimized MapReduce code, pig scripts and performance tuning and analysis
- Implemented advanced procedures like text analytics and processing, using the in-memory computing capabilities of Spark
- Exported the aggregated data into Oracle using Sqoop for reporting on the Tableau dashboard
- Involvement in design, development and testing phases of Software Development Life Cycle
- Performedinstallation, updates, patches and version upgrades when required for Hadoop
Environment: Hadoop, Map Reduce, HDFS, HBase, Hive, Pig, Java, SQL, Sqoop, Flume, Oozie, UNIX, Java, Maven, Eclipse
Data Integration Developer/Analyst
Confidential
Responsibilities:
- Involved in gathering specifications from Business User and designing of the process and set up time lines for entire process
- Extracting updated data Portals periodicallywith SAS/SQL
- Coordinating with different teams to make sure data is available on time
- Cleansing and validating data
- Creating data sets for analysis and report.
- Responsible to maintain previous month and previous financial year
- Responsible to validate data reports with historical data.
- Responsible for doing UAT before release to business users
- Performed documentation of the SAS Code for the better understanding of the program
- Involved in automation of reports in various formats like pdf, html and excel reports.
- Extensively used procedures like PROC SQL, PROC PRINT, and PROC SORTetc.
- Coded SAS programs with the use ofSAS jobs.
- Analyzed the data using SAS/STAT proceduresPROC FREQ,PROC MEANS,PROC
Environment: SAS, Oracle, SQL, Linux
Database Analyst/Developer
Confidential
Responsibilities:
- Work with Customer Analytics and key business stakeholders to prioritize the order in which disparate customer data sources are integrated into the customer data mart
- Provided analysis for senior management regarding ROI of promotions leading to shift marketing budget into the most effective channels
- Utilized advanced analytical methods in SAS and Microsoft Excel, including marketing mix models, to test the effectiveness of Hershey's promotional activities
- Employed forecasting models to understand underlying sales trends and expected future performance
- Monitor database performance, implement required changes
Environment: SAS, MS SQL Server, Excel, Windows
Java Developer
Confidential
Responsibilities:
- Used Object Oriented Programming and design.(OOP&OOD)
- Wrote stored procedures, complex queries using PL/SQL to extract data from the database, delete data and reload data on Oracle9i DB using the Toad tool.
- Developed both front-end and back-end of the product usingJava, J2EE, Ajax, JQuery, spring and Hibernate, and other technologies.
- Developed user interfaces using JSPs, HTML, CSS,JavaScript, jQuery, JSPCustomTags.
- Used Spring Core Annotations for Dependency Injection.
Environment: Java, J2EE, JSP, spring, Hibernate, Agile, Tomcat, Web Services, MySQL, Eclipse 3.5, SVN, Maven, JUnits, Hudson, JMS
