Bigdata/spark Developer Resume
Farmington, CT
SUMMARY:
- Overall 7+ years of IT experience including 4 years of Bigdata/Hadoop developer, 2 years of ETL Developer and 1 year of Java Developer.
- Hands on experience in Hadoop ecosystem including HDFS, Spark, MapReduce, Hive, Sqoop, Oozie, Flume, Kafka.
- Excellent knowledge on Hadoop Components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, YARN.
- Expertise in Java and Scala languages.
- Experience in Creating Hive tables and load the tables using Sqoop and processed data using Hive QL.
- Hands - on experience on RDD architecture, implementing Spark operations on RDD.
- Having knowledge on Spark Streaming to ingest data from multiple data sources into HDFS.
- Experience in data cleansing using Spark Functions.
- Familiar with Spark Context, Spark SQL, Data Frame and Pair RDD's.
- Hands-on experience in using relational databases like Oracle.
- Experience in importing and exporting the data using SQOOP from HDFS to Relational Database systems and vice-versa.
- Hands on Experience in creating tables, partitions and buckets in Hive.
- Extending Hive Core functionality by writing UDF’s for Data Analysis.
- Extensive experience in working with various distributions of Hadoop like enterprise versions of Cloudera (CDH5) and Hortonworks (HDP).
- Good Knowledge in Amazon Web Services.
- Experience working on EC2 (Elastic Compute Cloud) cluster instances, setup data buckets on S3 (Simple Storage Service), setting up EMR (Elastic MapReduce).
- Extensive programming experience in developing Java applications using Java, J2EE and JDBC.
- Experience in using different build tools like SBT and Maven.
- Good knowledge in NoSQL databases HBASE, MongoDB and Cassandra.
- Experience in using Producer and Consumer API’s of Apache Kafka.
- Expertise in Informatica client tools - PowerCenter Designer, Workflow Manager, Workflow Monitor and Repository Manager.
- Experience in Connected and Un-Connected Look up Transformations in the Designer of Informatica PowerCenter.
- Having good knowledge in using Unix and Linux.
- Experience in collection of JSON data into HDFS and processed the data using Hive and experienced in using Sequence files, AVRO file, Parquet file formats.
- Strong Knowledge on Python Language.
- Good in using version control like GITHUB and SVN.
- Managed the projects based on waterfall and Agile-Scrum Methods.
TECHNICAL SKILLS:
Big Data Ecosystems: Hadoop, MapReduce, HDFS, Spark, HBase, Zookeeper, Hive, Pig, Sqoop, Oozie, Flume, Kafka.
Programming Languages: Java, SQL, Scala, Python, HQL.
NoSQL Databases: HBase, Cassandra, MongoDB.
Databases: SQL Server, Oracle 8i/9i/10g.
Cloud: AWS.
Hadoop Distributions: Cloudera, Hortonworks.
Operating Systems: Microsoft Windows, LINUX, UNIX.
Office Tools: Microsoft Office '07, '10.
Development Tools: Eclipse, IntelliJ.
Build Tools: Maven, SBT.
Version Control Tools: GITHUB, SVN.
PROFESSIONAL EXPERIENCE:
Confidential - Farmington, CT
Bigdata/Spark Developer
Roles and Responsibilities:
- Worked on Cloudera distribution.
- Involved in extracting customer's data from various data sources to HDFS data lake which include data from relational RDBMS and csv files.
- Loaded and transformed large sets of structured and semi-structured data using Spark.
- Involved in working with Sqoop for loading the data from RDBMS to HDFS.
- Extensively used Spark Core, Spark SQL.
- Developed Spark applications Using Scala as per the Business requirements.
- Used Spark Data Frame Operations to perform required validations on the data.
- Responsible in performing sort, join, aggregations, filter, and other transformations on the datasets.
- Created Hive tables and working on them for data analysis to cope up with the requirements.
- Implemented Hive Partitioning and bucketing for data analytics.
- Analyzed the data by performing HQL, Spark SQL.
- Loaded the Cleaned Data into the hive tables and performed analytical functions based on requirements.
- Involved in creating views for the data security.
- Involved in the performance tuning of spark applications.
- Worked on Performance and Tuning operations in Hive.
- Involved in creating workflows to run Sqoop jobs monthly.
- Involved in Agile methodologies, daily Scrum meetings, Sprint planning.
- Experienced in using version control tools like GitHub to share the code snippet among the team members.
Environment: HDFS, Hive, Apache Sqoop, Spark, Scala, YARN, Agile Methodology, Cloudera, MySQL.
Confidential - Plano, TX
Bigdata/Spark Developer.
Roles and Responsibilities:
- Used Kafka to load data from different sources into Spark.
- Used spark to perform necessary transformations and actions on the data which comes from Kafka.
- Extensively used Spark Core, Spark SQL and Spark Streaming.
- Experienced in working with Spark SQL on different file formats like json and csv.
- Involved in accessing Hive tables in spark and analyzed the data using Spark SQL Queries.
- Responsible in performing sort, join, filter and other transformations on the datasets.
- Created Hive tables and working on them for data analysis to cope up with the requirements.
- Involved in Extending Spark functionality by writing custom UDFs in Scala.
- Implemented Hive Partitioning and bucketing for data analytics.
- Involved in working with Sqoop to export and import the data to RDBMS.
- Involved in the performance tuning of spark applications.
- Used Maven for building jar files of Spark programs and deployed to cluster.
- Implemented the workflows using Apache Oozie framework to automate tasks.
- Experience in importing and exporting data using Sqoop from HDFS to RDBMS and vice-versa.
- Involved in Agile methodologies, daily Scrum meetings.
- Experienced in using version control tools like GitHub to share the code snippet among the team members.
Environment: Cloudera, HDFS, Apache Spark, Apache Hive, Scala, Oozie, Apache Kafka, Apache Sqoop, Agile Methodology, Amazon S3.
Confidential
ETL Developer.
Roles and Responsibilities:
- Extensively used Informatica client tools - Source Analyzer, Warehouse designer, Mapping designer, Transformations, Informatica Repository Manager.
- Designed various mappings for extracting data from various sources involving flat files and relational tables.
- Involved in developing the mappings by using various transformations like source qualifier, sorter, aggregator, router, filter, lookup, expression etc.
- Created sessions, batches for incremental load into staging tables and scheduled them to run daily.
- Involved in preparing the Mapping design documents.
- Developed several reusable transformations and mapplets that were used in other mappings.
- Involved in preparing the Unit Test Cases for various mappings and workflows.
- Performance tuning of the process at the mapping level, session level, source level, and the target level.
- Developed workflows by using various tasks like command task, session, decision task and e- mail task.
- Worked with the Informatica Scheduler for scheduling the delta loads and master loads.
- Used Update Strategy to insert, delete, update and reject the items based on the requirement.
- Responsible for Production Support and Issue Resolutions using Session Logs, and Workflow Logs.
Environment: Informatica Power center 8.6.1, Oracle 10g, Windows XP, Unix Shell Scripts, SQL, PL/SQL, Flat files.
Confidential
Jr. Java Developer.
Roles and Responsibilities:
- Involved in requirement collection and analysis.
- Worked on developing front-end screens using JSP, Struts and HTML
- Involved in implementing persistent data management using JDBC.
- Participated in problem analysis and coding.
- Design and coding of screens involving complex calculations on various data windows accessing different tables on the oracle database.
- Developed screens for Patient Registration, Inventory of Medicines, Billing of Services and Asset Modules.
- Used JSF framework in developing user interfaces using JSF UI Components, validate Events and Listeners.
- Created several pieces of the JSF engine, including value bindings, bean discovery, method bindings, event generation and component binding.
- Involved in unit testing, integration testing, SOAP UI testing, smoke testing, system testing and user acceptance testing of the application.
- Wrote stored procedures, Database Triggers.
- Involved in debugging and troubleshooting related to production and environment issues
- Performed Unit testing.
Environment: JSP, Servlets, SQL, PL/SQL, WebSphere Application Server, Oracle 9i, JavaScript, windows XP, Unix shell Script, eclipse, MongoDB.
