Spark Scala Developer Resume
Tampa, FL
SUMMARY:
- Over 13+ years of IT experience with 4+years of hands on experience in Apache Hadoop ecosystem, Big Data analytics, and 9+ years as a IBM Mainframe and / Java professional experience in System Analysis, Development, Maintenance, Enhancement, Environment and Production Support for various projects assignments.
- Actively involved in the life cycle of the project, which includes implementing the Systems, Designing, Development, Testing and Documentation.
- Excellent ability to plan, organize, and prioritize my work and to meet on time the deadlines.
- Keen in building knowledge on emerging technologies in the Analytics, Information Management, Big data and related areas and in providing best business solutions
- Experienced in Hadoop environment technologies/Platforms like Pig, Hive, Sqoop.
- Consulting for Mainframe,Java and Big Data Technologies. Evaluation of new technologies in Big data, Analytics and NO Sql space
- Exclusive experience in Hadoop Ecosystem and its components like HDFS, Map Reduce, Yarn, Apache Pig, Hive, Sqoop, HBase, Flume, Kafka, Storm, Spark, Cassandra and Oozie.
- Involved in writing the Pig scripts to reduce the job execution time.
- Having experience with processing real time streamed data using Spark streaming, Spark SQL
- Strong Experience in working with Databases like DB2, SQL Server and proficiency in writing complex SQL, PL/SQL for creating tables, views, indexes, stored procedures and functions.
- Experience on Hadoop distributions like Cloudera, Hortonworks.
- Experienced in transporting, and processing real time event streaming using Kafka and Spark Streaming.
- Used XML Web Services using SOAP to transfer the data to application that is remote and global to different financial institutions.
- Experienced in Real time data ingestion into HBASE and HIVE using Spark.
- Experience with all stages of the SDLC and Agile Development model right from the requirement gathering to Deployment and production support.
- Also have experience in understanding of existing systems, maintenance and production support, on technologies such as Java, J2EE and various databases (DB2,Oracle, SQL Server).
- Excellent communication skills, interpersonal, hardworking and ability to proficiently communicate with all levels of the organization and work as a part of the team as well as independently.
- Hands - on experience working on Akka Streams
- Knowledge on FLUME, NO-SQL, MongoDB,Impala, Python, Talend, Tableau, Samza, Spark Ecosystem (ML Lib, Graph X)Data warehouse and BI technologies
- Implementing Unit Tests,System Integration Test(SIT) and User Acceptance Tests(UAT)
- Addressing users production queries/inquiries
- Hands-on experience leading a Team
- Highly Capable in learning things quickly and good at good time management.
- Focus on designing and delivering most optimum and critical business solutions for Big Data Technologies
TECHNICAL SKILLS:
Hardware: IBM 360/370/308X, 3090,4300,Personal Computers
Operating Systems: OS/390, TSO, MVS, Win 95/98, Windows NT, UNIX7, Windows XP/7/10, Linux(Cent OS, Ubuntu, Red Hat)
Languages: COBOL, Microfocus Cobol,PL/1, Assembler, JCL, SQL, PL/SQL, EZTRIEVE (Easytrieve), JAVA, J2EE, Scala, Python
Hadoop Distributions: Apache Hadoop, MapR, Hortonworks, Cloudera
Big Data Technologies: Hadoop, HDFS, MapReduce, YARN, Hive, Pig, Sqoop, HBase, Cassandra, Flume, Oozie, Spark, Kafka, Storm, Scala, Solr, Samza, Platfora
Web Technologies: J2EE(SERVLETS, JSP, JDBC), JavaScript, JQuery, AngularJs, NodeJS, Python
Web and Application Servers: JBoss 8.2, BEA Web Logic and Tomcat
Frameworks: Struts, Hadoop, Ext JS, Spring, Hibernet.
Java IDEs: Eclipse and My Eclipse, Net Beans.
Tracking Tools and Control: SVN, GIT
Databases: Oracle 11g, SQL Server 2008 R2, EDB(Postgre SQL),My SQL No SQL HBase, Cassandra, MongoDB
Data warehousing Tool: AbInitio
PROFESSIONAL EXPERIENCE:
Confidential,Tampa,FL
Spark Scala Developer
Responsibilities:- Installed and configured Hadoop Map Reduce, HDFS and developed multiple Map Reduce jobs in Java for data cleaning and preprocessing
- Collaborate with subject matter experts, various stakeholders and fellow developers to design, develop, implement and support data analytics.
- Created Hive tables, and loading and analyzing data using hive queries
- Having exposure to Teradata for processing the huge data
- Worked on debugging, performance tuning of Hive & Pig Jobs.
- Designed and implemented pig UDFs for evaluation, filtering, loading and storing of data
- Worked on Performance Tuning of Hadoop jobs by applying techniques such as MapSide Joins, Partitioning, Bucketing and using different file formats such as SequenceFile, RCFile, ORCFile
- Defined job work flows as per their dependencies in OOZIE.
- Used JAVA, J2EE application development skills with Object Oriented Analysis and extensively involved throughout Software Development Life Cycle (SDLC)
- Proactively monitored systems and services, architecture design and implementation of Hadoop deployment, configuration management, backup, and disaster recovery systems and procedures.
- Used Kafka to collect, aggregate, and store the web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
- Lead the team during data ingestion from external system.
- Load and transform large sets of structured, semi structured and unstructured data
- Supported Map Reduce Programs those are running on the cluster
- Importing and exporting data into HDFS and Hive using Sqoop, Spark Core and Spark SQL, Scala.
- Involved in loading data from UNIX file system to HDFS, configuring Hive and writing Hive UDFs
- Worked in converting Hive/SQL queries into Spark transformations using Spark RDDs and Scala.
- Developed Scala and SQL code to extract data from various databases.
- Developed Spark code using Scala and Spark-SQL for faster testing and data processing.
- Managed and reviewed log files
- Implemented partitioning, dynamic partitions and buckets in HIVE
- Analyzed vast data and discovered insights of the data using Tableau
Environment: HDP2.4, Hadoop, JDK1.6, Yarn, HDFS, Hive, Pig, Kafka, HBase, Spark, Scala, Sqoop, Akka Streams,Java, HTML, XML, SQL, J2EE, Eclipse, Oracle, Golden Gate, RC, ORC, Oozie, Teradata, Tableau
Confidential,Tampa, FL
Hadoop Developer
Responsibilities:- Used Sqoop to import data from RDBMS behind the JD Edwards Enterprise One ERP system to HDFS.
- Imported data (csv, plain text) from Amazon S3 using HDFS distributed copy command and s3n native filesystem URI.
- Wrote MapReduce clean-up programs, to ensure the data is in a consistent format, and to filter the unwanted records from the full data set.
- Setup Linux shell scripts and Oozie workflows to periodically import any incremental data generated, clean it, and add it into HDFS.
- Created Hive structures based on input data.
- Created Hive partitions based on the data and parameters specific to the functional requirements to optimize the performance.
- Wrote Hive queries and UDFs to process and transform data into different styles, calculate results for complex formulas etc.
- Automated the periodic analysis of data, and notified stakeholders via email also using Oozie workflows.
- Loaded data in HDFS to Spark RDD’s and performed various transformations, actions to process the data.
- Involved in migrating Hive queries into Spark transformations using Data frames, Spark SQL, SQLContext, and Scala.
- Utilized Kafka for loading streaming data and performed initial processing, real time analysis using Storm.
- Debugged various issues relating to connectivity with data sources, bad records in data, etc.
- Setup Hive ODBC connector, to make aggregated output data available to downstream BI applications.
- Closely worked with the Hadoop Admin to configure and maintain the cluster, the Hadoop Architect to design the HDFS folder structure for storing the inputs and outputs, the Data Analyst to agree on the output data format.
- Worked in a fast paced agile environment, delivered sprint goals, familiar with tools like JIRA Agile, Confluence, and Git Stash.
- Conducted knowledge transfer sessions to new members of the team about the application.
- Created company standard technical documentation with regards to architecture, design, test plans etc.
Environment: CDH4, Hadoop 2.0,HDFS, MapReduce, Yarn, Hive, Oozie, Sqoop, Oracle, Linux, Shell scripting,Java, Storm, Kafka, Eclipse, Amazon S3, JD Edwards Enterprise One, JIRA, Git Stash
Confidential,Tampa, FL
Mainframe and Java Developer
Responsibilities:- Build - Coding, Unit testing, Integration Testing, User Acceptance Testing and Production Support.
- Definition, development, and testing of the processes and programs necessary to extract data from the client's operational databases/files.
- Quality Activities, which includes performing quality review and quality check for all the deliverables made to client.
Environment: MVS, TSO/SPF-ISPF, COBOL, MQ Series, CA-7, JCL, VSAM, CICS, Core Java, Java Script, Web service, SQL Server 2008R2,Eclipse, TOAD, File-Aid, CHANGEMAN, Intertest
Confidential,Addison, TX
Responsibilities:- Requirements gathering and analysis -
- Build - Coding Testing - Unit testing, Integration Testing, User Acceptance Testing and Production Support
- Quality Activities, which includes performing quality review and quality check for all the deliverables made to client.
- Generating Reports by using Easytrieve for Sequential Files and DB2 tables.
- Definition, development, and testing of the processes and programs necessary to extract data from the client's operational databases/files.
- Significantly contributed in peer reviews test plans and test results, which are developed by the offshore team.
- Preparing weekly reports to be sent to the project lead.
Environment: OS/390, IBM 3090, TSO/SPF-ISPF, COBOL, JCL, DB2, SQL, CA-VIEW, VSAM, CICS, CHANGEMAN, Intertest, Startool.
Confidential
Responsibilities:- Analysis, Code building, Testing(Unit, SIT & UAT)
- Prepare programs specs, UTP & UTR.
- Involved in extensive review activities within the project for all kind of Service request from the customer.
- Providing permanent fixes for production problems and finally gave a complete resolution to the problem within SLA.
- Having daily meetings with the onsite team for discussing the project status.
- Maintaining Job documentation and Problem documentation.
Environment: MVS, TSO/SPF-ISPF, COBOL, CA-7, JCL, VSAM, CICS, File-Aid, CHANGEMAN, Intertest, VisionPlus 2.56.
Confidential
Responsibilities:- Design, Development, Implementation and support of the application.
- Thorough reviews and testing procedures of the developed components.
- Implemented the quality guidelines for the project Reviews, Defect prevention activities etc.
- Prepared Documentation related to all phases of project (Design, Test Plan, Review results, user manuals etc).
Environment: MVS, TSO/ ISPF, COBOL, JCL, DB2, SQL, VSAM, File-Aid, CHANGEMAN.
