I Hadoop/spark Scala Developer Resume
Irving, TX
SUMMARY:
- Around8 years of Software Development Experience in Java/J2EETechnology and Big data/Hadoopstack.
- Over 4+ years of experience in implementing complete Hadoop solutions, including data acquisition, data validation, data profiling, storage, transformation, analysis and integration with other frameworks to meet business needs using Big Data/Hadoop technology stack that includesHDFS, MapReduce, HBase, Hive, Pig, Impala, Oozie, Sqoop, Spark SQL, YARN/MRv1/MRv2, Flume(Web log processing),Kafka, Solr, Zookeeper, Strom, MongoDB, Cassandra. .
- Expertise in various Hadoop distributions like Amazon AWS Ec2/EMR, Cloudera, Hortonworks and MapR distributions.
- Experience in importing/exporting of data into/from Traditional Database like Teradata/Oracle RDBMS using Sqoop
- Experienced in getting streaming data into HDFs using Flume, memory channels, custom interceptors.
- Extensively worked in writing, fine tuning and profilingMapreducejobs for optimized performance.
- Extensive experience in implementing data analytical algorithms using Map reduce design patterns.
- Experience in implementing complex map reduce algorithms to perform joins on the Map side using distributed cache.
- Experience in writing Test cases, test classes usingMRUnit, Junit and Mockito.
- Extended Hive and Pig core functionality by writing Custom UDFs.
- Experienced in creating Hive internal/external tables, partitions, dynamic partitions, buckets to interact with data scientist to perform ad - hoc queries in structured data.
- Good understanding of various Hadoop file formats and compressions i.e. RCFile, Parquet, ORCFile, Avro, GZip and Snappy
- Experienced in handling ETL transformations using Pig Latin scripts, expressions, join operations and Custom UDF's for evaluation, filtering and storing data
- Expert inanalyzing real time queries using different NoSQL data bases including Cassandra and HBase.
- Expert in implementing advanced procedures like text analytics and processing using the in-memory computing capabilities like Apache Spark written in Scala.
- Experience in converting business process into RDD transformations using Apache Spark and Scala.
- Experience in Writing Producers/consumers and creating messaging centric applications using Apache Kafka.
- Have knowledge on Apache Storm to integrate with Apache Kafka for stream processing.
- Experience inintegrating Spark withSolrand Indexingwith Apache Solr.
- Experience in supporting data analysis projects using Elastic Map Reduce on the Amazon Web Services (AWS).
- Knowledge onSplunk UI to work at production support to perform log analysis.
- Expertise and Knowledge in usingjob scheduling and monitoring tools like Azkaban,OozieandZooKeeper.
- Expertise in writing Shell-Scripts, Cron Automation and Regular Expressions.
- Exposure in using Cloudera Manager and Apache Ambarifor monitoring cluster machines.
- Experience with variousSDLCmethodologies like Waterfall and Agile andObject Oriented Analysis and Design (OOAD).
- Strong hands on experience in development of Client/Server Applications using Java/J2EE, XML and MVC Frameworks like Spring, Struts and Hibernate.
- Strong knowledge in Web Services, SOA, ESB, SOAP, REST, WSDL, AXIS and JERSEY
- Experience in working with databases DB2, Oracle, MySQL. Extensive experience working on SQL, PL/SQL, using tables, triggers, views, packages and stored procedures
- Experienced in Object Oriented Analysis and Object Oriented Design using Unified Modeling Language (UML).
- Working knowledge of Web/Application Servers like JBoss, Apache Tomcat, IBM Web Sphere and Oracle Web Logic.
- Expertise in tools and utilities like Eclipse, TOAD for Oracle, Rational Rose (UML tool), WSAD, RAD, Ant, Maven.
- Strong knowledge of agile development methodologies, waterfall methodologies to minimize customer impact.
- Well-versed with all phases of testing such as unit testing with Junit, integration testing, Quality Assurance testing, System testing, UAT testing.
- Have good experience with both Windows,LINUX and UNIX platforms.
TECHNICAL SKILLS:
Programming Languages: C, Java, Scala
Distributed File Systems: Apache Hadoop HDFS
Hadoop Distributions: Amazon aws/EMR,Apache Cloudera, Hortonworks, and MapR
Hadoop Technologies: HDFS, MapReduce, Hive, Pig, Sqoop,Azkaban, Oozie, Zookeeper, Flumesparksql,and Apache Kafka
NoSQL data bases: Cassandra, Hbase
Relational Data Stores: Oracle, MySQL.
Search Platforms: Apache Solr
Inmemory/MPP/Search: Apache Spark, Apache Spark Streaming, Apache Storm
Operating Systems: Windows,UNIX,LINUX
Cloud Platforms: Amazon AWS, OpenStack.
Application Servers: JBoss, Tomcat, Web Logic, Web Sphere
Web Services: SOAP,REST, WSDL, JAXB, and JAXP
Frameworks: Hibernate, Spring, Struts, JMS, EJB
Web Technologies: HTML5, CSS3, AngularJS, JavaScript, JQuery, AJAX, Servlets, JSP,JSON, XML, XHTML, JSF
PROFESSIONAL EXPERIENCE:
Confidential, Irving TX
I Hadoop/Spark Scala Developer
Responsibilities:
- Responsible for design, development and delivery of data from operational systems and files into ODSs.
- Troubleshoot and develop on Hadoop technologies including HDFS, Hive, Spark, Impala and Hadoop ETL development.
- Working with Azure Data Platform components - Azure Data Lake, Data Factory, Data Management Gateway, Azure Storage Options, Azure SQL.
- Used Scala to convert Hive/SQL queries into RDD transformations in Apache Spark.
- Translate, load and present disparate data sets in multiple formats and multiple sources including Avro, text files.
- Responsible to develop the optimal performance strategy, and manage the technical metadata across all ETL jobs.
- Import the data from different sources like HDFS/Hive into Spark Data Frames and Data Sets with spark 2.0.
- Developed Apache spark jobs using Scala for faster data processing and used spark SQL for querying.
- Implemented the workflows using Apache Nifi framework to automate tasks.
- Created an Incremental Load with Change Detection Mechanism through Spark Scala Framework.
- Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
- Involved in creating Hive Tables, loading with data and writing Hive queries which will invoke and run Map Reduce jobs in the backend.
- Worked in designing and developing applications in Spark using Scala to compare the performance of Spark with Hive.
- Developed a Nifi Workflow to pick up the data from HDFS and drop it in SFTP server and send the email notifications.
- Working with BA’s, End users and architects to define and process requirements, build code efficiently and work in collaboration with the rest of the team for effective solutions.
- Deliver projects on-time and to specification with quality.
Confidential, Plano TX
Hadoop Developer/Supporting
Responsibilities:
- Supporting Data Platform built around Big Data Technologies on Hadoop, Spark, hive.
- Migration of Hadoop data from Data Registry to Cloud based distribution system Amazon S3.
- Identifies, diagnoses, and resolves functional and technical problems and business critical incidents through research and analysis of trends, root causes, and business impact by using HPSM Service Manager Tool.
- Highlights functionality issues to developers.
- Assists in the translation of solutions into technical requirements.
- Raise defect reports to the development team for code amendment.
- Analysis is involved in the full life cycle of an application and part of an agile development process.
- Ability to interact, develop, engineer, and communicate collaboratively at the highest technical levels with clients, development teams.
- Scheduling and monitoring the production jobs by using Control-M and Arow Tools.
- Implementing the Code changes, Quotes, based on the developer requirement.
- Active participation in Kafka, EMR rehydration whenever new security updates are released by amazon distribution.
- Creating Instant S3 Buckets and EMR clusters for the developers to process the data.
Confidential, Austin TX
Spark Scala Developer
Responsibilities:
- Actively participated with the development team to meet the specific customer requirements and proposed effective Hadoop solutions
- Developed Spark scripts by using Scala shell commands as per the requirement
- Processing the schema oriented data using Scala and Spark.
- Developed and designed automate process using shell scripting for data movement and generating Confidential format VCF Files.
- Developed Scala scripts, UDFFs using both Data frames/SQL/Data sets and RDD in Spark 1.6 for Data Aggregation, queries and writing data back into files saving in to the hadoop.
- Used Scala collection framework to store and developed Enrichments to process the complex consumer information.
- Used Scala functional programming concepts to develop business logic .
- Exported and Imported data from SQL databases like IBMDB2 ,by developing spark scala scripts and into Big Data Lake
- Structured data was ingested onto the data lake using Sqoop jobs and scheduled using Oozie workflow from the IBM DB2 data sources for the incremental data.
- Involved in converting business transformations into Spark Data Frames and RDDs using Scala.
- Involved in integrating hive queries into spark environment using SparkSql and Spark Scala.
- Data was processed using spark such as aggregating, calculating the statistical values by using different transformations and actions.
- Involved in Migrating Abinitio ETL tool to Spark Scala Frame Work to Extract the Custom Graphs.
- Computing the complex logics and controlling the Data flow through In-memory process tool Apache Spark.
- Developed Spark code using Scala and Spark-SQL for faster testing and processing of data.
- Experienced in configuring work flows using Oozie.
- Used various spark Transformations and Actions for cleansing the input data
- Mentored analyst and test team for writing Hive queries.
- Involved in deploying multi module applications using Maven and Jenkins.
- Experienced in working in agile environment and on-site/offshore co-ordination.
Environment: Hadoop Framework,Hive, Sqoop, Oozie, Maven, Jenkins,Java(JDK1.6), Scala, UNIX Shell Scripting, Oracle 11g/12g,IBM Db2.
Confidential, L os Angeles CA
Data Engineering
Responsibilities:
- Design and develop ELT data pipeline using Spark App to fetch data from Legacy system and thirdparty API, social media sites.
- Perform data analytics and load data to Amazon s3/datalake/Spark cluster.
- Write and buildAzkabanworkflow jobs to automate the process.
- Develop spark sql tables & queries to perform adhoc data analytics for analyst team.
- Deploy components using Maven Build system and Docker images
- Involved in deploying multi module Azkaban applications using Maven
- Processed the source data to structured data and store in NoSQL database.
- For automating deployment process developed Shell Scripts.
- Played an important in migrating jobs from spark 0.9 to 1.4 to 1.6.
- Experience in production on-call support, worked on incidents.
- Implemented Spark RDD transformations to map business analysis and apply actions on top of transformations.
- Design and develop DMA( Confidential Movies anywhere) dashboard for BI analyst team.
- Monitoring spark clusters .
- Experience in Agile methodology.
- Hands on experience in using Zeeplin and Databricks tools.
- Worked with file formats TEXT, AVRO, PARQUET and SEQUENCE files.
- Developed ETL Frame Works.
- Implemented UDFS, UDAFS, UDTFS in java for hive to process the data that can’t be performed using Hive inbuilt functions
- Involved in creating Hive Internal/External tables, loading with data and troubleshoot with Hive jobs.
- Created partitioned tables in Hive for best performance and faster querying.
- Programming and/or Scripting experience with Python, Spark, etc.
- Experience in building and debugging PySpark jobs with Python programming.
- Involved in integrating hive queries into spark environment using SparkSql.
- Participated in Brain stroming,Sprint Review,Retrosepective Meetings
- Participate in requirement gathering and analysis phase of the project in documenting the business requirements by conducting meetings with various business users/Product owners.
Environment: Languages/Technologies: Java(JDK1.6 and higher), Azkaban, Spark, Spark Sql, Presto, Hive, Apache Crunch, Elastic Search, Spring boot.Special Software: Azkaban, Eclipse, GIT Repository, Amazon S3, Amazon AWS Ec2/EMR, Amazon EMR Spark cluster,Hadoop Framework, Sqoop, Azkaban, Maven, UNIX Shell Scripting
Confidential, Coppell TX
Hadoop Developer
Responsibilities:
- Actively participated with the development team to meet the specific customer requirements and proposed effective Hadoop solutions
- Structured data was ingested onto the data lake using Sqoop jobs and scheduled using Oozie workflow from the RDBMS data sources for the incremental data.
- Streaming Data (Time Series Data) was ingested into the data lake using Flume.
- Implemented custom interceptors for flume to filter data and defined channel selectors to multiplex the data into different sinks.
- Developed Map Reduce programs using Java programming language that are implemented on the Hadoop cluster.
- Used Avro data serialization system with Avro tools to handle Avro data files using Map reduce programs.
- Implemented Data Validation using map reduce programs to remove un-necessary records before move data into Hive tables.
- Implemented UDFS, UDAFS, UDTFS in java for hive to process the data that can’t be performed using Hive inbuilt functions
- Implemented optimized map side joins to get data from different data sources, cleaning data .
- Designed and implemented custom writable, custom input formats, custom partitions and custom comparators.
- Involved in creating Hive Internal/External tables, loading with data and troubleshoot with Hive jobs.
- Used the RegEx, JSON and Avro SerDe’s for serialization and de-serialization packaged with Hive to parse the contents of streamed log data and implemented Hive custom UDF’s.
- Experienced in Using Hive ORC formats for better columnar format, compression and processing.
- Wrote pig scripts for advanced analytics on the data for recommendations.
- Gained profound knowledge of using Cassandra.
- Processed the source data to structured data and store in NoSQL database Cassandra.
- Worked with the performance testing team to optimize the Cassandra cluster by making certain changes in Cassandra yamlconfiguration file and some Linux OS configurations
- Involved in converting business transformations into Spark RDDs using Scala.
- Involved in integrating hive queries into spark environment using SparkSql.
- Computing the complex logics and controlling the Data flow through In-memory process tool Apache Spark.
- Along with the Infrastructure team, involved in design and developed Kafka and Storm based data pipeline.
- Implemented messaging system for different data sources using apache Kafka and configuring High level consumers for online and off-line processing.
- Experienced in configuring work flows using Oozie.
- Involved in deploying multi module applications using Maven and Jenkins.
- Experienced in working in agile environment and on-site/offshore co-ordination.
Environment: Hadoop Framework, MapReduce, Hive, Sqoop, Pig, HBase, Cassandra, Apache Kafka, Storm, Flume, Oozie, Maven, Jenkins,Java(JDK1.6), UNIX Shell Scripting, Oracle 11g/12g.
Confidential, Minneapolis MN
Java/Hadoop Developer
Responsibilities:
- Designed and implemented Java engine and API to perform direct calls from front-end JavaScript (ExtJS) to server-side Java methods (ExtDirect).
- Developed interfaces and their implementation classes to communicate with the mid-tier (services) using JMS. Technically, it is a 3-tier client server application, where GUI tier interacts with Java middle-tier custom library and queries an Oracle 10g database using Hibernate.
- Implemented Restful services for Account Summary and workable list etc for Reports Decouple support .
- Developed and consumed SOAP Web services using JBoss ESB framework
- Designed and developed Java batch programs in Spring Batch.
- Involved in write database schema Through Hibernate.
- Installed and configured Hadoop MapReduce, HDFS, Developed multiple MapReduce jobs in java for data cleaning and pre-processing.
- Experienced in managing and reviewing Hadoop log files
- Developed Sqoop scripts to import export data from relational sources and handled incremental loading on the customer, transaction data by date.
- Tuning of MapReduce configurations to optimize the run time of jobs.
- For automating the cluster installation developed Shell Scripts.
- Wrote the shell scripts to monitor the health check of Hadoop daemon services and respond accordingly to any warning or failure conditions.
- Involved in loading data from UNIX file system to HDFS.
- Created java operators to process data using DAG streams and load data to HDFS.
- Developed custom Input Formats to implement custom record readers for different datasets.
- Used Pig as ETL tool to do transformations, event joins, filter and some pre-aggregations.
- Automated all jobs for pulling data from FTP server to load data into Hive tables using Oozie workflow.
- Experienced in using Java Rest API to perform CURD operations on HBase data.
- Create & deploy web services using REST framework.
- Developed suit of Unit Test Cases for Mapper, Reducer and Driver classes using MR Testing library.
- Developed Unit test cases using Junit and MRUnit testing frameworks.
Environment: Hadoop, MapReduce, HDFS, Hbase, Hive, Java, SQL, Cloudera Manager, Sqoop, Flume, Oozie, Shell Scripts, Java JDK 1.6, Eclipse.
Confidential
Java Developer
Responsibilities:
- Coded front end components using HTML, JavaScript and jQuery, Back End components using Java, spring, Hibernate, Services Oriented components using Restful and SOAP based web services, and Rules based components using JBoss Drools.
- Developed presentation layer using JSP, HTML and CSS and JQuery.
- Developed JSP custom tags for front end.
- Written Java script code for Input Validation.
- Used Apache CXF open source tool to generate java stubs form WSDL.
- Developed and consumed SOAP Web services using JBoss ESB framework
- Developed the Web Services Client using SOAP, WSDL description to verify the credit history of the new customer to provide a connection.
- Developed RESTful Web services within ESB framework and used Content based Routing to route to ESB's
- Designed and developed Java batch programs in Spring Batch.
- Test Driven development is done by maintaining the Junit and FlexUnit test cases throughout the application.
- Developed stand-alone Java batch applications with springand Hibernate.
- Involved in write database schema Through Hibernate.
- Designed and developed DAO layer with Hibernate standards, to access data from IBM DB2.
- Developed the UI panels using JSF, XHTML, CSS, DOJO and JQuery.
Environment: Java 6 - JDK 1.6, JEE, Spring 3.1 framework, Spring Model View Controller (MVC), Java Server Pages (JSP) 2.0, Servlets 3.0, JDBC4.0, AJAX, Web services, Rest full, JSON, Java Beans, JQuery, JavaScript, Oracle 10g, JUnit, HTML Unit, XSLT, HTML/DHTML.
