We provide IT Staff Augmentation Services!

Bigdata/spark Developer Resume

2.00/5 (Submit Your Rating)

Farmington, CT

SUMMARY:

  • Overall 7+ years of IT experience including 4 years of Bigdata/Hadoop developer, 2 years of ETL Developer and 1 year of Java Developer.
  • Hands on experience in Hadoop ecosystem including HDFS, Spark, MapReduce, Hive, Sqoop, Oozie, Flume, Kafka.
  • Excellent knowledge on Hadoop Components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, YARN.
  • Expertise in Java and Scala languages.
  • Experience in Creating Hive tables and load the tables using Sqoop and processed data using Hive QL.
  • Hands - on experience on RDD architecture, implementing Spark operations on RDD.
  • Having knowledge on Spark Streaming to ingest data from multiple data sources into HDFS.
  • Experience in data cleansing using Spark Functions.
  • Familiar with Spark Context, Spark SQL, Data Frame and Pair RDD's.
  • Hands-on experience in using relational databases like Oracle.
  • Experience in importing and exporting the data using SQOOP from HDFS to Relational Database systems and vice-versa.
  • Hands on Experience in creating tables, partitions and buckets in Hive.
  • Extending Hive Core functionality by writing UDF’s for Data Analysis.
  • Extensive experience in working with various distributions of Hadoop like enterprise versions of Cloudera (CDH5) and Hortonworks (HDP).
  • Good Knowledge in Amazon Web Services.
  • Experience working on EC2 (Elastic Compute Cloud) cluster instances, setup data buckets on S3 (Simple Storage Service), setting up EMR (Elastic MapReduce).
  • Extensive programming experience in developing Java applications using Java, J2EE and JDBC.
  • Experience in using different build tools like SBT and Maven.
  • Good knowledge in NoSQL databases HBASE, MongoDB and Cassandra.
  • Experience in using Producer and Consumer API’s of Apache Kafka.
  • Expertise in Informatica client tools - PowerCenter Designer, Workflow Manager, Workflow Monitor and Repository Manager.
  • Experience in Connected and Un-Connected Look up Transformations in the Designer of Informatica PowerCenter.
  • Having good knowledge in using Unix and Linux.
  • Experience in collection of JSON data into HDFS and processed the data using Hive and experienced in using Sequence files, AVRO file, Parquet file formats.
  • Strong Knowledge on Python Language.
  • Good in using version control like GITHUB and SVN.
  • Managed the projects based on waterfall and Agile-Scrum Methods.

TECHNICAL SKILLS:

Big Data Ecosystems: Hadoop, MapReduce, HDFS, Spark, HBase, Zookeeper, Hive, Pig, Sqoop, Oozie, Flume, Kafka.

Programming Languages: Java, SQL, Scala, Python, HQL.

NoSQL Databases: HBase, Cassandra, MongoDB.

Databases: SQL Server, Oracle 8i/9i/10g.

Cloud: AWS.

Hadoop Distributions: Cloudera, Hortonworks.

Operating Systems: Microsoft Windows, LINUX, UNIX.

Office Tools: Microsoft Office '07, '10.

Development Tools: Eclipse, IntelliJ.

Build Tools: Maven, SBT.

Version Control Tools: GITHUB, SVN.

PROFESSIONAL EXPERIENCE:

Confidential - Farmington, CT

Bigdata/Spark Developer

Roles and Responsibilities:

  • Worked on Cloudera distribution.
  • Involved in extracting customer's data from various data sources to HDFS data lake which include data from relational RDBMS and csv files.
  • Loaded and transformed large sets of structured and semi-structured data using Spark.
  • Involved in working with Sqoop for loading the data from RDBMS to HDFS.
  • Extensively used Spark Core, Spark SQL.
  • Developed Spark applications Using Scala as per the Business requirements.
  • Used Spark Data Frame Operations to perform required validations on the data.
  • Responsible in performing sort, join, aggregations, filter, and other transformations on the datasets.
  • Created Hive tables and working on them for data analysis to cope up with the requirements.
  • Implemented Hive Partitioning and bucketing for data analytics.
  • Analyzed the data by performing HQL, Spark SQL.
  • Loaded the Cleaned Data into the hive tables and performed analytical functions based on requirements.
  • Involved in creating views for the data security.
  • Involved in the performance tuning of spark applications.
  • Worked on Performance and Tuning operations in Hive.
  • Involved in creating workflows to run Sqoop jobs monthly.
  • Involved in Agile methodologies, daily Scrum meetings, Sprint planning.
  • Experienced in using version control tools like GitHub to share the code snippet among the team members.

Environment: HDFS, Hive, Apache Sqoop, Spark, Scala, YARN, Agile Methodology, Cloudera, MySQL.

Confidential - Plano, TX

Bigdata/Spark Developer.

Roles and Responsibilities:

  • Used Kafka to load data from different sources into Spark.
  • Used spark to perform necessary transformations and actions on the data which comes from Kafka.
  • Extensively used Spark Core, Spark SQL and Spark Streaming.
  • Experienced in working with Spark SQL on different file formats like json and csv.
  • Involved in accessing Hive tables in spark and analyzed the data using Spark SQL Queries.
  • Responsible in performing sort, join, filter and other transformations on the datasets.
  • Created Hive tables and working on them for data analysis to cope up with the requirements.
  • Involved in Extending Spark functionality by writing custom UDFs in Scala.
  • Implemented Hive Partitioning and bucketing for data analytics.
  • Involved in working with Sqoop to export and import the data to RDBMS.
  • Involved in the performance tuning of spark applications.
  • Used Maven for building jar files of Spark programs and deployed to cluster.
  • Implemented the workflows using Apache Oozie framework to automate tasks.
  • Experience in importing and exporting data using Sqoop from HDFS to RDBMS and vice-versa.
  • Involved in Agile methodologies, daily Scrum meetings.
  • Experienced in using version control tools like GitHub to share the code snippet among the team members.

Environment: Cloudera, HDFS, Apache Spark, Apache Hive, Scala, Oozie, Apache Kafka, Apache Sqoop, Agile Methodology, Amazon S3.

Confidential

ETL Developer.

Roles and Responsibilities:

  • Extensively used Informatica client tools - Source Analyzer, Warehouse designer, Mapping designer, Transformations, Informatica Repository Manager.
  • Designed various mappings for extracting data from various sources involving flat files and relational tables.
  • Involved in developing the mappings by using various transformations like source qualifier, sorter, aggregator, router, filter, lookup, expression etc.
  • Created sessions, batches for incremental load into staging tables and scheduled them to run daily.
  • Involved in preparing the Mapping design documents.
  • Developed several reusable transformations and mapplets that were used in other mappings.
  • Involved in preparing the Unit Test Cases for various mappings and workflows.
  • Performance tuning of the process at the mapping level, session level, source level, and the target level.
  • Developed workflows by using various tasks like command task, session, decision task and e- mail task.
  • Worked with the Informatica Scheduler for scheduling the delta loads and master loads.
  • Used Update Strategy to insert, delete, update and reject the items based on the requirement.
  • Responsible for Production Support and Issue Resolutions using Session Logs, and Workflow Logs.

Environment: Informatica Power center 8.6.1, Oracle 10g, Windows XP, Unix Shell Scripts, SQL, PL/SQL, Flat files.

Confidential

Jr. Java Developer.

Roles and Responsibilities:

  • Involved in requirement collection and analysis.
  • Worked on developing front-end screens using JSP, Struts and HTML
  • Involved in implementing persistent data management using JDBC.
  • Participated in problem analysis and coding.
  • Design and coding of screens involving complex calculations on various data windows accessing different tables on the oracle database.
  • Developed screens for Patient Registration, Inventory of Medicines, Billing of Services and Asset Modules.
  • Used JSF framework in developing user interfaces using JSF UI Components, validate Events and Listeners.
  • Created several pieces of the JSF engine, including value bindings, bean discovery, method bindings, event generation and component binding.
  • Involved in unit testing, integration testing, SOAP UI testing, smoke testing, system testing and user acceptance testing of the application.
  • Wrote stored procedures, Database Triggers.
  • Involved in debugging and troubleshooting related to production and environment issues
  • Performed Unit testing.

Environment: JSP, Servlets, SQL, PL/SQL, WebSphere Application Server, Oracle 9i, JavaScript, windows XP, Unix shell Script, eclipse, MongoDB.

We'd love your feedback!