We provide IT Staff Augmentation Services!

Sr. Spark And Hadoop Developer Resume

3.00/5 (Submit Your Rating)

EXPERIENCE SUMMARY:

  • 9+ years of IT experience this includes 1 year of C++ and Java in the field of software development and maintenance, 5 years of Teradata Experience in the field of data warehouse application development and 3+ year of Spark and Hadoop experience in the field of software development and maintenance.
  • Experienced with Spark Context, Spark - SQL, Data Frame, Pair RDD's, Spark YARN.
  • Experience in using Spark-SQL with various data sources like JSON, Parquet and Hive.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark and Scala.
  • Experience in writing spark scripts for SCD type 1 & 2 incremental loads using scala.
  • Experienced in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself.
  • Good knowledge on Hadoop Ecosystem components like Spark, HDFS, Map Reduce, Hive, Pig, Impala, Sqoop and Oozie.
  • Experience in generating Hadoop performance metrics using Cloudera Manager that portrays the overall cluster health status on weekly basis for senior management at the bank.
  • Created Hive table with Static and Dynamic partitions and bucketing.
  • Experience in using DML statements to perform different operations on Hive Tables.
  • Ingested large volume of RDBMS transactional data source into HDFS using Sqoop import & analysed with Hive.
  • Developed Sqoop import jobs with incremental load to populate Hive External tables.
  • Developed User Defined Functions (UDF’s) in Java as and when necessary to use in HIVE queries.
  • Developed Oozie workflows for Map Reduce, Hive, Pig and Sqoop actions.
  • Having good knowledge on Kafka and Spark Streaming
  • Commendable knowledge / experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems (RDBMS) and vice-versa.
  • Hadoop Shell commands, Writing Map reduce Programs, Verifying the Hadoop Log Files.
  • Proficient in producing production support platform reporting and metrics for 'W' data warehousing..
  • Good knowledge on Unix commands and shell scripting.
  • Having Good knowledge on PL/SQL.
  • Followed Agile Methodologies while working on the project.
  • Performance tuning on complex queries with efficiency of PI/SI indexes, Join Index, PPI, Using Explain analysing the data distribution among AMPs and index usage, collect statistics, definition of indexes, revision of correlated sub queries, etc.
  • MAPR CERTIFIED HADOOP DEVELOPER (MCHD) & Teradata 12 Certified Solutions Developer
  • Good knowledge in relational database like Oracle and Teradata and programming languages like C, C++, Java and some of the scripting languages like Unix Shell scripting and JavaScript.
  • Have hands on experience in Onsite-Offshore model projects.
  • Quick learner with short learning curve
  • Good communication, interpersonal, time-management skills as well as a great team player.
  • Excellent analytical abilities, initiative, creative, managing skills in different technologies.

TECHNOLOGY:

Operating System: UNIX, CentOS Linux, Windows 98/2000/NT/XP, MS DOS

Big Data Ecosystem: Spark, Hadoop, Map Reduce, HDFS, Hive, Pig, Sqoop, Impala, Oozie

RDBMS: Teradata and Oracle

Languages: C, C++, JAVA, SCALA, SQL, PL/SQL, UNIX Shell.

Other tools: CITRIX, Shell Scripts, K-Shell Scripts and VI-editor, Eclipse IDE, WinSCP, Github, Putty and UC4

PROFESSIONAL EXPERIENCE:

Confidential

Sr. Spark and Hadoop Developer

Responsibilities:

  • Understanding the business requirement and preparation of Business Requirement specification. Study the source system and map the requirement with it.
  • Create mapping documents using Data Modelling and Data Lineage documents
  • Develop the scripts for loading data from source systems to target by applying necessary transformations using Spark and Scala.
  • Review the code and run unit tests and integration tests
  • Deploy jar and run your application in a spark cluster
  • QA & UAT support and bug fixing.

Tools: Scala, Hive, Spark core, Spark SQL and UC4

Confidential

Spark Developer

Responsibilities:

  • Implemented Spark and Spark SQL for faster testing and processing of data.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs.
  • Optimizing of existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames and Pair RDD's.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself.
  • Words counting for every 10 seconds that provided through TCP Sockets.

Tools: Hive, Spark core, Spark SQL and Spark Streaming

Confidential

Hadoop Developer

Responsibilities:

  • Extract the data from Confidential and store in Dataset and use SQL and Hive context to do the processing of data and store the result in RDBMS.
  • Processing Confidential file using HIVE and Spark.
  • Creating different table using HIVE and process the data and stored the result in different hive table.
  • Export the resulting table to an RDBMS using SQOOP.
  • Compared hive results with spark SQL and found the best results through spark with low latency.

Tools: Eclipse, Hive, Sqoop, Spark and MySQL

Confidential

Hadoop Developer

Responsibilities:

  • Created Hive table with Static and Dynamic partitions and bucketing.
  • Experience in using DML statements to perform different operations on Hive Tables.
  • Ingested large volume of RDBMS transactional data source into HDFS using Sqoop import &analyzed with Pig/Hive.
  • Developed Sqoop import jobs with incremental load to populate Hive External tables.
  • Created Pig scripts to transform raw data from several data sources into forming baseline data.
  • Developed User Defined Functions (UDF’s) in Java as and when necessary to use in HIVE queries.
  • Developed Oozie workflows for Map Reduce, Hive, Pig and Sqoop actions.

Tools: Eclipse, HDFS, Map Reduce, Hive, Pig, Sqoop, MySQL

Confidential

Technical Analyst

Responsibilities:

  • Review post week’s Health and status of the Production Support Platforms: Production(VA8), RHUB(VA20), ADP(TX16),AHP(VA9) &HTX(TX18) - CPU usage, AWT usage, query response times with query volume, system trend report, system resource usage, CPU usage, Utility usage, Archival usage, Utility exhaustion, Top CPU users. Mainframe MIPS &CPU usage in LPAR. CPU Usage on Informatica and IIS server.
  • Finding average query response times in different boxes like VA8, VA20 and TX16.
  • Experience in generating reports using PDCR, DBQL and DBC tables for senior management at the bank. These reports are for various requirements like assessing CPU/IO/space usage for different applications.
  • Impact CPU and Heavy Hitter identification and suggesting the load teams for tuning bad performed queries.
  • Experience in generating performance metrics that portrays the overall system health on daily, weekly and monthly basis for senior management at the bank.
  • Co-ordinate with the ETL load team and other development teams for SLA compliance/reporting.
  • Build SharePoint sites for easy data management.
  • Co-ordinate meetings for software and hardware upgrades.
  • Prepare pre and post upgrade performance reports.
  • Experience in generating following Hadoop performance metrics using Cloudera Manager that portrays the overall cluster health status on weekly basis for senior management at the bank.
  • CPU and Memory Utilization for all edge nodes and Data nodes
  • Disk Space Utilization on all mount points for all the edge nodes
  • Disk Space and Memory Utilization on Name Node
  • Edge Node Disk Utilization by application
  • Job tracker Memory used
  • Average map &reduce task running
  • RPC average processing time Remote procedure calls
  • HDFS Cluster Disk usage by applications
  • Healthy task tracker
  • Block distribution across all PROD data nodes

Operating System: Windows 7, Linux

Solution Environment: Teradata and SQL

Tools: Teradata RDBMS 12.0/13.10/14.10, BTEQ, FASTLOAD, MLOAD, TPMUP, Teradata SQL Assistant, HDFS, Map Reduce, Sqoop, Hive, Pig, Oozie, Cloudera Manager and WHEM Portal.

Confidential

Upgrade Tester

Responsibilities:

  • Involved in Upgrade testing of the developed objects.
  • Create Project Node: - This is a onetime task per baseline. We need to create a new project node.
  • Create Admin Unit: - This is onetime task per baseline. We need to create admin unit using devop tools.
  • Create a Kit: - This is repeat task for each upgrade.
  • Upgrade on VM: - This is repeat task for each upgrade.
  • Post Upgrade Verification: - This is repeat task for each upgrade.
  • Address all the critical issues which arise as part of development of the product and minimize the bugs in the code.
  • Perform upgrade testing to the product for every build to ensure that errors are not crept in to the code.
  • Identifying the root cause for all the upgrade failures
  • Perform Upgrade test for all projects and making sure that promote code to the Release only on successful upgrade.
  • If upgrade fails due to project code then fix the issue and re-perform the upgrade testing.
  • If upgrade fails due to code outside of project code, coordinate with the corresponding code owner and the upgrade team, have the other team fix the code, and promote all the code together into the same baseline thus making sure the baseline upgrade does not fail.
  • Interacting with the onsite team / clients.
  • Impact analysis within and outside the system
  • Finding the best solution for customer raised PRs(problem report) and do the Upgrade testing.

Operating System: Windows XP/NT, UNIX

Solution Environment: C++ and Java

Tools: C++, Java, Eclipse IDE, Clear case, Putty, MS-Visio, UNIX Shell.

We'd love your feedback!