We provide IT Staff Augmentation Services!

Spark Developer(hadoop) Resume

3.00/5 (Submit Your Rating)

RockvillE

SUMMARY

  • Over 10+ years of experience in IT industry which includes 4 years in BigData ecosystem related technologies.
  • Around 7 years of IT experience in developing BI Applications Using Informatica, Oracle and Unix
  • About 4 years of Data analytics experience on designing and implementing complete end - to- end Big Data/Hadoop Infrastructure solutions using HDFS, PIG, HIVE, HBase, Sqoop, Flume, Oozie, MapReduce,
  • 3+ years’ experience in Scala, Apache Spark Core, Spark SQL, Spark Streaming
  • 2+ years’ experience wif developing Hadoop applications using Java
  • 3 years of experience in NoSQL databases Hbase, Mongodb
  • 2 years of experience in Oracle Business Intelligence Enterprise Edition (11.x)
  • Around 2 years of Experience in Python Scripting.
  • Around 1-year experience wif Programming and Statistical Testing and Modeling wif R Language
  • Around 1-year experience wif Tableau
  • 1-year experience on AWS - S3 for storage, EC2 and EMR for processing/analysis
  • 1-year Experience wif Hadoop Administration
  • 2-years’ Experience wif Kafka Streaming
  • Excellent understanding / noledge of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
  • Hands on experience in installing, configuring, and using Hadoop ecosystem components like Hadoop MapReduce, HDFS, HBase, Hive, Sqoop, Pig, Zookeeper, Oozie and Flume.
  • Experience in Managing scalable Hadoop clusters including Cluster designing, provisioning, custom configurations, monitoring and maintaining using different hadoop distributions: Cloudera CDH, MapR, Apache Hadoop
  • Experience in data analysis using HiveQL, Pig Latin, HBase and custom Map Reduce programs in Java.
  • Experience in Oozie and workflow scheduler to manage Hadoop jobs by Direct Acyclic Graph (DAG) of actions wif control flows.
  • Good understanding of NoSQL databases and hands on work experience in writing applications on NoSQL databases like HBase and MongoDB.
  • Experienced in working wif Amazon Web Services (AWS) using EC2 for computing and S3 as storage mechanism.
  • Strong Experience of Data Warehousing ETL concepts using Informatica Power Center, OLAP, OLTP and AutoSys.
  • Experience in Job management using Fair scheduler and Developed job processing scripts using Oozie workflow.
  • Populated HDFS wif huge amounts of data using Apache Kafka and Flume
  • Designing and creating Hive external tables using shared meta-store instead of derby wif partitioning, dynamic partitioning and buckets.
  • Well experienced in data transformation using custom MapReduce, Hive and Pig scripts for different types of file formats Text File, Sequence File, Avro, ORC, and PARQUET
  • Experience in analyzing data using HiveQL, Pig Latin, HBase and custom programs in Java,Python
  • Extending Hive and Pig core functionality by writing custom UDFs.
  • Experience in managing and reviewing Hadoop log files.
  • Hands on experience in application development using Oracle(SQL,PL/sqL),Pyhon,shell scripting.
  • Knowledge of various NoSQL storage technologies (Key-Value, Column-Family, Document)
  • Experience in managing lifecycle of MongoDB database including database sizing, deployment automation, monitoring and tuning
  • Experience database backups and test recoverability regularly and overall performance of teh mongodb
  • Experience wif development, deployment, and support of large-scale distributed applications in a mission-critical production environment.
  • Experience wif Building stream-processing systems wif Kafka
  • Expert in creating SQL Queries, PL/SQL Packages, Function, stored procedures, triggers, and cursors, created database objects like tables, views, sequences, synonyms, indexes using Oracle tools like SQL*Plus, SQL Developer and Toad.
  • Proficient in advance features of Oracle 11g for PL/SQL programming like Using Records and Collections, Bulk Bind, Ref. Cursors, Nested tables and Dynamic SQL, Oracle Advanced queues.
  • Extensively worked on ETL using Informatica - Power Center
  • Designed complex Mappings and expertise in performance tuning.
  • Automation and scheduling of UNIX shell scripts and Informatica sessions and batches using Autosys.
  • Experienced in Tuning Informatica Mappings to identify and remove processing bottlenecks
  • Experience in performance tuning of Oracle BI Repository & Dashboards / Reports, by implementing Aggregate tables, Indexes and managing Cache.
  • Expertise in generating customized reports using OBIEE
  • Good Experience in UNIX Shell Script, Perl Script for automation of tasks for file loading, Job Scheduling.
  • Good Experience in Agile methodology
  • Good experience wif Build and configuration tools Autosys, Harvest,Udeploy,SBT Eclipse, Maven, Jenkins, SVN, JIRA, Client AML, Control M, Git tools
  • Domain Experience in Financial, Telecom, Healthcare
  • Highly motivated team player wif good communication skills and excellent problem-solving abilities. Would be willing to work independently or as part of a team.

TECHNICAL SKILLS

Bigdata Technologies: Hadoop 1.x and 2.x, HDFS,Hive,MapReduce,Pig, SqoopFlume,Zookeeper,Yarn,Kafka,oozie,flume,Spark(1.x and 2.x) Core, Spark SQL,Spark Streaming,kafka

Languages: SQL, PL/SQL, Core Java,Scala

Web Technologies: Java Script, Json HTML, XML.AWS

Oracle Tools: SQL Developer, SQL* Plus, Eclipse, Toad,OBIEE(10g,11g)

Data integration: Informatica 8.x,9.x

Databases: Oracle 8i/9.2/10g/11g,Mysql,Teradata

No SQL Databases: HBase,Mongodb

Operating System: Windows 7/XP/NT/Vista, MS-DOS, Linux and UNIX.

Scripting: Unix Shell script, Perl, Python, Java Script

PROFESSIONAL EXPERIENCE

Confidential, Rockville

Spark Developer(Hadoop)

Responsibilities:

  • Involved in start to end process of Hadoop jobs dat used various technologies such as Sqoop, PIG, Hive, Spark and Shell scripts (for scheduling of jobs ) Extracted and loaded data into Data Lake environment
  • Solid understanding and experience in applying and implementing machine learning algorithms and concepts such as: Classification and Regression, Resampling statistics and bootstrapping using R language
  • Experience in working wif Hadoop 2.x version and Spark 2.x (Python and Scala).
  • Extending HIVE and PIG core functionality by using custom User Defined Function's (UDF), User Defined Table-Generating Functions (UDTF) and User Defined Aggregating Functions (UDAF) in Java.
  • Developed simple and complex MapReduce programs in Java for Data Analysis on different data formats
  • Implemented business logic by writing Hive UDFs in Java.
  • Experience in Hadoop Production support tasks by analyzing teh Application and cluster logs.
  • Implemented Partitioning, Dynamic Partitions, and Buckets in Hive on Avro files to meet teh business requirements.
  • Expertise in designing and deployment of Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, Oozie, Zookeeper, SQOOP, flume, Spark, Kafka, Hbase wif MapR Distribution.
  • Assist in upgrading, configuration and maintenance of various Hadoop infrastructures like Pig, Hive, and Hbase.
  • Used Spark for interactive queries, processing of streaming data and integration wif NoSQL database for huge volume of data.
  • Developed Scala scripts using both Data frames/SQL and RDD/MapReduce in Spark 1.x/2.x for Data Aggregation, queries and writing data back into OLTP system through Sqoop.
  • Used Spark-Streaming APIs to perform necessary transformations and actions on teh fly for building teh common learner data model which gets teh data from Kafka in near real time and Persists into Hbase.
  • Performed advanced procedures like text analytics and processing, using teh in-memory computing capabilities of Spark using Scala.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself.
  • Worked on migrating Map Reduce programs into Spark transformations using Spark and Scala.
  • Coordinate wif Administration team to enhance Spark Jobs performance by analyzing them.
  • Implementeddesign patternsin Scala for teh Spark application.
  • Developed quality code adhering to Scala coding Standards and best practices.
  • Used Spark API over MapR Hadoop YARN to perform analytics on data in Hive.
  • Explored wif teh Spark improving teh performance and optimization of teh existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Pair RDD's, Spark YARN.
  • Worked on Loading log data into HDFS using Flume, Kafka and performing ETL integrations
  • Used Reporting Tool Tableau to connect wif Hive for generating daily reports of data.
  • Collaborated wif teh infrastructure, network, database, application and BA teams to ensure data quality
  • Worked wif different File Formats like TEXTFILE, SEQUENCE FILE, AVROFILE, ORC, and PARQUET for Hive querying and processing
  • Developed Spark code using scala and Spark-SQL/Streaming for faster testing and processing of data.
  • Import teh data from different sources like HDFS/Hbase into Spark RDD.
  • Load teh data into Spark RDD and do in memory data Computation to generate teh Output response.
  • Experience in Oozie and workflow scheduler to manage Hadoop jobs by Direct Acyclic Graph (DAG) of actions wif control flows.
  • Configure Oozie workflow to run multiple Hive and Pig jobs which run independently wif time and data availability.
  • Collected and aggregated large amounts of log data using Apache Flume and staging data in HDFS for further analysis
  • Developed a data pipeline using Kafka and Storm to store data into HDFS.
  • Developed business specific Custom UDF's in Hive, Pig.
  • Responsible for developingPythonwrapper scripts which will extract specific date range using Sqoop by passing custom properties required for teh workflow
  • Skilled in using collections inPythonfor manipulating and looping through different user defined objects.
  • Worked wif different kind of compression techniques like LZO, GZip, Snappy.
  • Worked on various configurations of Oozie bundles for Orchestrating Pig, Hive, Spark, Sqoop

Environment: RedHat Linux, MapR,Scala,PythonR language, Python HDFS Hive,Pig,Sqoop, Flume, Oozie, Hbase, Spark Core, Spark SQL, Spark streaming, Kafka,Tableau

Confidential, Cleavland,Ohio

Lead Developer(Hadoop)

Responsibilities:

  • Developed data pipeline using Flume, Sqoop, Pig and map reduce to ingest various sources data into HDFS for analysis.
  • Developed job flows in Oozie to automate teh workflow for extraction of data from warehouses and weblogs.
  • Used Pig as ETL tool to do transformations, event joins, filter bot traffic and some pre-aggregations before storing teh data onto HDFS
  • Optimizing Map reduce code,pig scripts, user interface analysis, performance tuning and analysis .
  • Used Hive to analyze teh partitioned and bucketed data and compute various metrics for reporting on teh dashboard.
  • Developed Pig Latin scripts to extract and filter relevant data from teh web server output files to load into HDFS.
  • Developed simple and complex MapReduce programs in Java for Data Analysis on different data formats
  • Implemented business logic by writing Pig UDF's in Java and used various UDFs from Piggybanks and other sources.
  • Responsible for building scalable distributed data solutions using Hadoop
  • Installed and configured Hive, Pig, Sqoop, Flume and Oozie on teh Hadoop cluster
  • Developed Hive queries and Pig scripts to customize teh large data sets into JSON.
  • Involved in loading JSON datasets into MongoDB and validating teh data using Mongo shell.
  • Loaded teh aggregated data into MongoDB for reporting on teh dashboard.
  • Worked on MongoDB schema/document modeling, querying, indexing and tuning
  • Developed Simple to complex Map-reduce Jobs using Hive and Pig
  • Optimized Map-Reduce Jobs to use HDFS efficiently by using various compression mechanisms
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and extracted teh data from oracle into HDFS using Sqoop
  • Created and maintained Technical documentation for all teh tasks performed like executing Pig scripts and Hive queries
  • Created Hbase tables to store various data formats of PII data coming from different portfolios.
  • Installed Oozie workflow engine to run multiple Hive and pig jobs.
  • Worked on tuning teh performance Pig queries.
  • ImplementedSparkusing Scala andSparkSQL for faster testing and processing of data.
  • Involved in converting Hive/SQL queries intoSparktransformations usingSparkRDDs, Scala.
  • Created Spark RDD for data-centric task processing using Scala
  • ImplementedSparkusing Scala andSparkSQL for faster testing and processing of data.
  • Hadoop usingSparkContext,Spark-SQL, Data Frame, Pair RDD's &SparkYARN.
  • Used Informatica for HADOOP for loading data to and from HDFS and HIVE tables
  • Worked on Slowly Changing Dimensions both Type1 and Type 2.
  • Designing ETL processes using Informatica to load data from Flat Files, Oracle and Excel files to target Oracle Data Warehouse database.
  • Developed mappings in Informatica to load teh data from various sources into teh Data Warehouse, using different transformations like Joiner, Aggregator, Update Strategy, Rank, Router, Lookup, Sequence Generator, Filter, Sorter, Source Qualifier.
  • Designed workflows wif many sessions wif decision, assignment task, event wait, and event raise tasks, used Autosys scheduler to schedule jobs

Environment: RedHat Linux,Cloudera HDFS, MapReduce, Hive, Pig, Sqoop, Flume, Zookeeper, Oozie, Hbase,Mongodb,Informatica 9.x,OBIEE 11.3,Spark Core,Spark SQL,Spark streaming

Confidential

Developer(Hadoop,Informatica,oracle pl/sql)

Responsibilities:

  • Installed and configured Hadoop MapReduce, HDFS, developed multiple Map Reduce jobs in java for data cleaning and pre-processing.
  • Developed multiple MapReducejobs in python for data cleaning and preprocessing.
  • Designed Oozie workflows.
  • Installed and configured Hive and also written Hive UDFs.
  • Implemented CDH3 Hadoop cluster.
  • Installing cluster, monitoring/administration of cluster recovery, capacity planning, and slots configuration.
  • Created HBase tables to store variable data formats of PII data coming from different portfolios.
  • Implemented best income logic using Pig scripts.
  • Exported teh analyzed data to teh relational databases using Sqoop for visualization and to generate reports for teh BI team.
  • Developed Shell and Python scripts to automate and provide Control flow to Pig scripts.
  • Worked on MongoDB schema/document modeling, querying, indexing and tuning and Involved in loading JSON datasets into MongoDB and validating teh data using Mongo shell.
  • Supported in setting up QA environment and updating configurations for implementing scripts wif Pig and Sqoop.
  • Writing Hadoop MR programs to get teh logs and feed into Cassandra for Analytics purpose
  • Building, packaging and deploying teh code to teh Hadoop servers.
  • Unix Scripting to manage teh Hadoop Operation stuffs.
  • Wrote Stored Procedures, Functions, Packages and triggers using PL/SQL to implement business rules and processes.
  • Extensive testing ETL experience using Informatica 9.x (Power Center/ Power Mart) (Designer, Workflow Manager, Workflow Monitor and Server Manager)
  • Worked on Informatica Power Center tools- Designer, Repository Manager, Workflow Manager, and Workflow Monitor.
  • Used advanced SQL like analytical functions, aggregate functions for mathematical and statistical calculations.
  • Optimized SQL used in reports to improve performance dramatically.
  • Tuned and optimized teh complex SQL queries.
  • Worked wif Business users to gather requirements for developing new Reports or changes in teh existing Reports.

Environment: Hadoop, MapReduce, HDFS, Hive, Python, SQL,, PIG, Sqoop, CentOS,Cloudera.Oracle 10g,11g, Autosys,Shellscripting,MongoDB.OBIEE11g,Informatica 9.x

Confidential

Developer(Informatica,oracle pl/sql)

Responsibilities:

  • Perform Data Mapping and develop ETL Specification documents
  • Develop ETLs using PL/SQL in Oracle 10g & 11g to extract, transform and load data from OLTP into Warehouse
  • Designing ETL processes using Informatica to load data from Flat Files, Oracle and Excel files to target Oracle Data Warehouse database.
  • Developed mappings in Informatica to load teh data from various sources into teh Data Warehouse, using different transformations like Joiner, Aggregator, Update Strategy, Rank, Router, Lookup, Sequence Generator, Filter, Sorter, Source Qualifier
  • Using Workflow Manager to create Sessions and scheduled them to run Confidential specified time wif required frequency
  • Utilized PL/SQL nested tables for conditional trafficking of data wifin ETL process.
  • Messaging & filtering of data by developing reusable PL/SQL functions
  • Performance tuning using Oracle Hints and Result caching where appropriate.
  • Utilized PL/SQL bulk collect feature to optimize teh ETL performance.
  • Develop BASH shell scripts to set up batch jobs on Unix Solaris 10 server.
  • Document application from a technical maintenance point of view and host a noledge transfer session.
  • Maintain and Enhance Oracle PL/SQL batch process for patient level data collected in a clinical trial and reporting system
  • Load teh data from MS Excel to Oracle Table and Oracle table to MS Excel.
  • Loaded teh data using teh SQL loader, Imports & UTL files based on file formats.
  • Debugging Production issues using Toad 9.7 debugger
  • Used SQL Trace and TKProf for analyzing performance issues
  • Perform Root Cause Analysis to identify and deploy bug fixes

Environment: Oracle 10g, UNIX, Windows, Informatica 8.6

Confidential

Developer(oracle pl/sql,Unix)

Responsibilities:

  • Involved in teh Extraction, Transformation and loading of teh data from various sources into teh dimensions and teh fact tables in teh Data Warehouse.
  • Created reusable Transformations and Mapplets and used them in various mappings.
  • Involved in extensive performance tuning by determining bottlenecks Confidential various points like targets, sources, mappings, sessions or system. This led to better session performance.
  • Created Informatica mappings wif PL/SQL procedures to build business rules to load data.
  • Most of teh transformations were used like teh Source qualifier, Aggregators, Connected & Unconnected lookups, Filters & Sequence.
  • Coding and development of packages in packages, procedures, cursors, tables, views and function as per teh business requirements.
  • Supported QA and resolved teh defects raised by QA.
  • Created Cursors and Ref cursors as a part of teh procedure to retrieve teh selected data.
  • Fine Tuned procedures for teh maximum efficiency in various schemas across databases using Oracle Hints, Explain plan and Trace sessions.
  • Written complex SQLs using joins, sub queries and correlated sub queries.

Environment: Oracle 10g, PL /SQL, UNIX, Windows, Shell, DPL(Data Presentation Language), PLSQL & Perl.

Confidential

Developer(Informatica,oracle pl/sql)

Responsibilities:

  • Interacted wif business community and gathered requirements based on changing needs and incorporated identified factors into Informatica mappings to build Data Marts.
  • Extensively worked on Power Center Designer to develop mappings using several transformations such as Filter, Joiner, Lookup, Rank, Sequence Generator, Aggregator and Expression transformations.
  • Implemented Type 1 and Type 2 Slowly Changing Dimensions.
  • Created parameter files and used mapping parameters and variables for incremental loading of data.
  • Involved in performance Tuning of Transformations, mappings, and Sessions for better performance.
  • Designing Tables, Constraints, Views, and Indexes etc. in coordination wif teh application development
  • Developed stored procedures/function on request for enhancement of business logic
  • Developed Job Scheduler scripts for data migration using UNIX Shell scripting.
  • Building teh different applications under Telegence and Light speed Billing Application
  • Involved in automation of build activity of teh project using Shell Script.
  • Building Emergency Fixes for different applications
  • Developed Data migration scripts using UTL FILE package

Environment: Oracle 10g, PL /SQL, UNIX, Windows, Shell, CVS, Power Builder, Change Tracker (CT), Mercury Quality Center,Informatica 8.x

We'd love your feedback!