Sr. Spark/scala Developer Resume
St Louis, MO
SUMMARY:
- Software developer with 8+ years of technical expertise in all phases of Software development cycle (SDLC), in various Industrial sectors like Banking, Financial, Auto Insurance, Health Care expertizing in Big data analyzing Frame works and Java/J2EE technologies
- 4+ years of industrial experience in Big Data analytics, Data manipulation, using Hadoop Eco system tools Map - Reduce, HDFS, Yarn/MRv2, Pig, Hive, HDFS, HBase, Spark, Kafka, Flume, Sqoop, Flume, Oozie, Avro, Sqoop, AWS, Spark integration with Cassandra, Avro, Solr and Zookeeper.
- Extensive experience in developing applications that perform Data Processing tasks using Teradata, Oracle, SQL Server and MySQL database.
- Hands on expertise in working and designing of Row keys & Schema Design with NOSQL databases like Mongo DB 3.0.1, Cassandra, H Base and DynamoDB (AWS) .
- Extensively worked on Spark using Scala on cluster for computational (analytics), installed it on top of Hadoop performed advanced analytical application by making use of Spark with Hive and SQL/Oracle.
- Excellent Programming skills at a higher level of abstraction using Scala, Java and Python.
- Experience in using D-Streams, Accumulator, Broadcast variables, RDD caching for Spark Streaming.
- Hands on experience in developing Spark applications using Spark tools like RDD transformations, Spark core, Spark MLlib, Spark Streaming and Spark SQL.
- Strong experience and knowledge of real time data analytics using Spark Streaming, Kafka and Flume.
- Working knowledge of Amazon’s Elastic Cloud Compute (EC2) infrastructure for computational tasks and Simple Storage Service (S3) as Storage mechanism.
- Running of Apache Hadoop, CDH and Map-R distros, dubbed Elastic MapReduce (EMR) on (EC2).
- Expertise in developing Pig Latin scripts and using Hive Query Language.
- Developed Customized UDFs and UDAF’s in java to extend HIVE and Pig core functionality.
- Created Hive tables to store structured data into HDFS and processed it using HiveQL.
- Worked on GUI Based Hive Interaction tools like Hue, Karmasphere for querying the data.
- Experience in validating, cleansing data using Pig statements and hands-on experience in developing Pig MACROS.
- Working knowledge in installing and maintaining Cassandra by configuring the Cassandra.yaml file as per the business requirement and performed reads/writes using Java JDBC connectivity.
- Experience in creating MongoDB clusters and hands on experience with complex MongoDB aggregate functions and mapping
- Experience in writing Complex SQL queries, PL/SQL, Views, Stored procedure, triggers, etc.
- Experience in OLTP and OLAP design, development, testing and support of enterprise Data warehouses.
- Written multiple MapReduce Jobs using Java API, Pig and Hive for data extraction, transformation and aggregation from multiple file formats including Parquet, Avro, XML, JSON, CSV, ORCFILE and other compressed file formats Codecs like gZip, Snappy, Lzo.
- Good experience in optimizing Map Reduce algorithms using Mappers, Reducers, combiners and partitioner’s to deliver the best results for the large datasets.
- Good knowledge on build tools like Maven, Log4j and Ant.
- Experienced in migrating data from different sources using PUB-SUB model in Redis, and Kafka producers, consumers and preprocess data using Storm topologies.
- Had competency in using Chef, Puppet and Ansible configuration and automation tools. Configured and administered CI tools like Jenkins, Hudson Bambino for automated builds.
- Hands on experience in using various Hadoop distros (Cloudera (CDH 4/CDH 5), Hortonworks, Map-R, IBM Big Insights, Apache and Amazon EMR Hadoop distributions.
- Knowledge in installation, configuration, supporting and managing Hadoop Clusters using Apache, Cloudera (CDH3, CDH4) distributions and on Amazon web services (AWS).
- Experienced in writing Ad Hoc queries using Cloudera Impala, also used Impala analytical functions.
- Experience in using Flume to load log files into HDFS and Oozie for data scrubbing and process
- Have the working about Knowledge about Splunk architecture and various components (Indexer, forwarder, search head, deployment server)
- In depth understanding/knowledge of Hadoop Architecture and various components such as HDFS, MapReduce Programming Paradigm, High Availability and YARN architecture.
- Proficient in developing, deploying and managing the Solr from development to production.
- Experience in importing data using Sqoop, SFTP from various sources like RDMS, Teradata, Mainframes, Oracle, Netezza to HDFS and performed transformations on it using Hive, Pig and Spark.
- Used various Project Management services like JIRA for tracking issues, bugs related to code and GitHub for various code reviews and Worked on various version control tools like CVS, GIT, SVN.
- Experienced in working with monitoring tools to check status of cluster using Cloudera manager, Ambari, Ganglia and Nagios.
- Used Docker in deploying various micro services as Containers, given networking routes, written several Docker Files as per the current requirement and used them in various environments.
- Hands-on knowledge in Core Java concepts like Exceptions, Collections, Data-structures, I/O. Multi-threading, Serialization and deserialization of streaming applications.
- Experience in maintaining an Apache Tomcat MYSQL, LDAP, LAMP, Web service environment.
- Ability to work with Onsite and Offshore Teams.
- Done Clustering, regression and Classification using Machine learning libraries Mahout, MLlib(Spark).
- Hands on experience in App Development using Java, Hadoop, RDBMS and Linux Shell scripting.
- Good experience with use-case development, with Software methodologies like Agile and Waterfall.
- Good understanding of all aspects of Testing such as Unit, Regression, Agile, White & Black-box.
- Proven ability to manage all stages of project development Strong Problem Solving and Analytical skills and abilities to make Balanced & Independent Decisions.
TECHNICAL SKILLS:
Big Data Ecosystem: HDFS, MapReduce, Pig, Hive, Spark, YARN, Kafka, Flume, Sqoop, Solr, Impala, Oozie, ZooKeeper, Spark, Ambari, Mahout, MongoDB, Cassandra, Avro, Parquet and Snappy.
Hadoop Distributions: Cloudera (CDH3, CDH4, and CDH5), Hortonworks, MapR and Apache
Languages: Java, Python, Scala, SQL, JavaScript and C/C++
No SQL Databases: Cassandra, MongoDB, HBase and Amazon Dynamodb.
Java Technologies: JSE, Servlets, JavaBeans, JSP, JDBC, JNDI, AJAX, EJB and struts
XML Technologies: XML, XSD, DTD, JAXP (SAX, DOM), JAXB
Web Design Tools: HTML, DHTML, AJAX, JavaScript, JQuery and CSS, AngularJs and JSON
Development / Build Tools: Eclipse, Jenkins, Git, Ant, Maven, IntelliJ, JUNIT and log4J.
App/Web servers: WebSphere, WebLogic, JBoss and Tomcat
DB Languages: MySQL, PL/SQL, PostgreSQL and Oracle
RDBMS: Teradata, Oracle 10g,11i, MS SQL Server, MySQL and DB2
Operating systems: UNIX, Red Hat LINUX, Mac os and Windows Variants
Testing: Hadoop MRUNIT Testing, Hive Testing, Quality Center (QC)
ETL Tools: Talend, Informatica, Pentaho, Ab Initio
PROFESSIONAL EXPERIENCE:
Confidential, St Louis, MO
Sr. Spark/Scala Developer
Responsibilities:
- Developed Spark Applications using Scala, Java and Implemented Apache Spark data processing project to handle data from various RDBMS, No-SQL DB’s and Streaming sources.
- Worked with Spark using Scala for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Spark MLlib, Data Frame, Pair RDD's, Spark YARN.
- Used Spark Streaming APIs to perform transformations and actions on the fly for building common learner data model which gets the data from Kafka in Near real time and persist it to Cassandra .
- Deployed KAFKA connect in standalone and distributed mode creating Docker containers using DOCKER, had modified the Docker Files as per the current platform.
- Developed Kafka consumer's API in Scala for consuming data from Kafka topics and Consumed XML messages using Kafka and processed xml file using Spark Streaming to capture UI updates .
- Used Apache Kafka functionalities to aggregate web log data from multiple servers and make them available in Downstream systems for Data analysis and engineering type of roles.
- Developed Preprocessing job using Spark Data frames to flatten Json documents to flat file.
- Loaded D-Stream data into Spark RDD and do in memory data Computation to generate Output response.
- Worked and learned a great deal from AWS Cloud services like EC2, S3, EBS, RDS and VPC.
- Migrated an existing on-premises, Physical-DC application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, Experienced in Maintaining the Hadoop cluster on AWS EMR.
- Imported data from AWS S3 to Spark RDD, performed transformations and actions on RDD's using Scala.
- Worked with ELASTIC MAPREDUCE and setup Hadoop environment on AWS EC2 Instances.
- Implemented S3 by creating a bucket and passing the event information to Lambda .
- Had Good understanding of Cassandra architecture and designed Columnar families in Cassandra, Ingested data from RDBMS, performed data transformations, and then exported the transformed data to Cassandra, also used Spark DataStax Cassandra Connector to load data to and from Cassandra.
- Experienced in Creating data-models for Client’s transactional logs, analyzed the data from Casandra tables for quick searching, sorting and grouping using the Cassandra Query Language (CQL).
- Used Hive QL to analyze the partitioned and bucketed data, Executed Hive queries on Parquet tables stored in Hive to perform data analysis to meet the business specification logic.
- Experience in using Avro, Parquet, RCFile and JSON file formats, developed UDFs in Hive and Pig.
- Worked with Log4j framework for logging debug, info & error data.
- Deployed Microservices like Spark, Cassandra on Maraton DC/OS, Hadoop clusters using Docker.
- Used Amazon DynamoDB to gather and track the event-based metrics.
- Developed Sqoop and Kafka Jobs to load data from RDBMS, External Systems into HDFS and HIVE.
- Developed Oozie coordinators to schedule Pig and Hive scripts to create Data pipelines.
- Extensive experience in using Microservices, Marathon (prod, test, dev) environments and Docker .
- Experienced in configuring the yaml files for Spark, MongoDB , Cassandra and deployed in Docker for connecting to the several Microservices .
- Setting up and worked on Kerberos authentication principals to establish secure network communication on cluster and testing of HDFS, Hive, Pig and MapReduce to access cluster for new users.
- Modified ANT Scripts to build JAR's, Class, WAR and EAR files, used Jenkins for Continuous integration and Bamboo Agents for deployment using Marathon Make system.
- Generated various kinds of reports, created Dash boards using Splunk, using the indexed logs into Splunk integrated appwatcher feature to Splunk and monitored the health of micro services using Splunk.
- Used Jira for bug tracking, Bit Bucket to check-in and checkout code changes, worked with QA, BI teams to ensure data quality and availability and responsible for generating actionable insights from complex data.
- Worked in Agile Methodology environment in delivering agreed user stories on time for every Sprint.
Environment : Docker- Containerized world, Hadoop 2.7.0, YARN, HDFS, Spark 2.2, Sqoop 1.99.7, Hive 2.1.1, Flume 1.7.0, Oozie 4.2.0 HDP 2.5, Pig 0.16.0, Kafka 0.9.0, Hbase 1.1.2, Zookeeper 3.4.8, AWS-EC2, S3, Talend Studio, Jenkins 2.0, Splunk, Cassandra, Scala, Java 8, Superputty, Scala IDE, Bit Bucket .
Confidential, Freeport, ME
Spark/Hadoop Developer
Responsibilities:
- Worked on migrating MapReduce programs into Spark transformations using Spark and Scala, initially done using python (PySpark).
- Developed Spark jobs using Scala on top of Yarn/MRv2 for interactive and Batch Analysis.
- Experienced in querying data using SparkSQL on top of Spark engine for faster data sets processing.
- Built and deployed the scalable distributed Hadoop cluster running on Hortonworks Data Platform (HDP 2.6) in the custom stack on the physical Data Center which is on-premises.
- Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS and store in databases such as HBase .
- Involved in creating Data Lake by extracting Application, customer’s data from various data sources to HDFS/HBase which include data from Databases, and logging data from servers and Micro services.
- Developing programs for Spark streaming which takes the data from Kafka and pushes into different sources like S3, HDFS and to local file system if necessary based on config.
- Involved in loading data from REST Endpoints to Kafka producers and transferring the data to Kafka Brokers
- Partitioning Data Streams using Kafka. Designed and configured Kafka Cluster to accommodate heavy throughput.
- Created, altered and deleted topics ( Kafka Queues ) when required with varying Performance tuning using Partitioning, bucketing of IMPALA tables and Convert the data into relational format to load into HBase/HDFS .
- Experience on Kafka and Spark integration for real time data processing
- Experienced with batch processing of data sources using Apache Spark and Elastic search.
- Used Impala connectivity from the User Interface (UI) and query the results using ImpalaQL.
- Loading the data from the different Data sources like ( Teradata, DB2, Oracle and flat files ) into HDFS using Sqoop and load into Hive tables , which are partitioned.
- Implemented Partitioning, Bucketing in Hive for better organization of the data.
- Developed Oozie Workflows for daily incremental loads, which gets data from Teradata and then imported into Hive tables.
- Developed bash Scripts to bring the log files from FTP server and then processing it to load into Hivetables , Bash Scripts are scheduled using Resource Manager Scheduler.
- Implemented Kafka event Log producer to produce the logs into Kafka Topic which are utilized by ELK (Elastic Search, Log Stash, Kibana) stack to analyze the logs produced by the Hadoop Cluster
- Created HBase tables, used HBase sinks and loaded data into them to perform analytics using Tableau.
- Used Elastic Search & MongoDB for storing and querying the offers and non-offers data.
- Experience in creating tables, dropping and altered at run time without blocking updates and queries using HBase and Hive, configured pig, written Pig Latin scripts to convert data from Text file to Avro format.
- Created Partitioned Hive tables and worked on them using HiveQL and Loading Data into HBase using Bulk Load and Non-bulk load.
- Involved in Data Engineering kinds of roles, where my role is to Ingest data to HDFS, and make necessary transformations as per the requirement.
- Worked on continuous Integration tools Jenkins and automated jar files at end of day.
- Worked on No-SQL database, MongoDB for POC purpose in storing images and URIs.
- Involved in loading data from UNIX file system to HDFS. Proficiency in Unix/Linux shell commands
- Worked with Tableau and Integrated Hive, Tableau Desktop reports and published to Tableau Server.
- Developed Unix shell scripts to load large number of files into HDFS from Linux File System.
- Worked in Agile development environment having KANBAN methodology. Actively involved in daily Scrum and other design related meetings.
- Used OOZIE Operational Services for batch processing and scheduling workflows dynamically.
Environment: Spark, Spark-Streaming, Spark SQL, Horton Works, HDFS, HBase, Hive, Pig, Apache Kafka, Sqoop, Java (JDK SE 7), Mongo-DB, Redis, Scala, Python, PySpark, Shell/Bash scripting, Linux, MySQL Oracle Enterprise DB, SOLR, Jenkins, Eclipse, Oracle, Git, Oozie, Tableau, MySQL, Soap, NIFI, Cassandra and Agile Methodologies.
Confidential, Naperville, Illinois
Big Data Developer
Responsibilities:
- Experienced in migrating and transforming of large sets of Structured, semi structured and Unstructured RAW data from RDMS, Oracle DB, Tera Data through Sqoop and placed in HDFS for further processing.-
- Written multiple Map Reduce programs in Java for data extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV and other codec file formats.
- Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
- Created multiple Hive tables, running hive queries in those data, implemented Partitioning, Dynamic Partitioning and Buckets in Hive for efficient data access
- Experienced in running batch processes using Pig Latin Scripts and developed Pig UDFs for data manipulation according to Business Requirements, Analysed the data by running Pig scripts.
- Got good experience with NOSQL database like MongoDB, HBase.
- Worked on MongoDB for distributed storage and processing.
- Used PIG to perform data validation on the data ingested using Sqoop and Flume and the cleansed data set is pushed into MongoDB.
- Designed and implemented the MongoDB schema, wrote services to store and retrieve user data from the MongoDB for the application on devices.
- Used Mongoose API to access the MongoDB from NodeJS.
- Written Java program to retrieve data from HDFS and providing it to REST Services.
- Implemented Sqoop for large data transfers from RDMS to HDFS/HBase/Hive.
- Implemented partitioning, bucketing in Hive for better organization of the data .
- Involved in using HCATALOG to access Hive table metadata from Map Reduce or Pig code
- Installed, configured and maintained Flume, Hive, Pig, Sqoop and Oozie on the Hadoop cluster.
- Created multiple Hive tables, running hive queries in those data, implemented Partitioning, Dynamic Partitioning and Buckets in Hive for efficient data access
- Experienced in running batch processes using Pig Latin Scripts and developed Pig UDFs for data manipulation according to Business Requirements
- Hands on experience in Developing optimal strategies for distributing the web log data over the cluster, importing and exporting of stored web log data into HDFS and Hive using Scoop.
- Developed several REST web services which produces both XML and JSON to perform tasks, leveraged by both web and mobile applications.
- Developed Unit test cases for Hadoop M-R jobs and driver classes with MR Testing library.
- Continuously monitored and managed the Hadoop cluster using Cloudera manager and Web UI .
- Designed the logical and physical data model, generated DDL scripts, and wrote DML scripts for Oracle 10g database
- Followed Agile Methodology for entire project and supported testing teams.
Environment: Hadoop, HDFS, Hive, Map Reduce, SOLR, Impala, MySQL, Oracle, Sqoop, Kafka, Spark, SQL Talend, Yarn, Pig, MongoDB, Oozie, Linux-Ubuntu, Scala, Ab Initio, Tableau, Maven, Jenkins, Java (JDK 1.6), Cloudera, JUnit, Agile methodologies
Confidential, Orrville, Ohio
Hadoop Developer
Responsibilities:
- Worked on analysing Hadoop cluster using different big data analytic tools including Pig, Hive, and MapReduce.
- Migrated existing SQL queries to HiveQL queries and made several POCs to move physical data to big data analytical platforms.
- Optimized Map/Reduce Jobs to use HDFS efficiently by using various compression mechanisms and got experience in managing and reviewing Hadoop Log files.
- Developed data pipeline using Pig and Hive from Teradata, DB2 data sources. These pipelines had customized UDF'S to extend the ETL functionality.
- Involved in the Installation and configuration of Flume, Hive, Pig, Sqoop and Oozie on the Hadoop Cluster using the Cloudera’s CDH distribution.
- Supported Data Analysts in running MapReduce Programs and Worked on processing unstructured data using Pig and Hive.
- Experienced with working on Avro Data files using Avro Serialization system
- Experienced in running Hadoop streaming jobs to process terabytes of xml format data.
- Collected and aggregated large amounts of weblog data from different sources using ApacheFlume and stored the data into HDFS for analysis.
- Experience in creating integration between Hive and HBase for effective usage and performed MR Unit testing for the Map Reduce jobs.
- Designed logical, Physical data model, generated DDL scripts, and wrote DML scripts for Oracle 9i database
Environment: Apache Hadoop, MapReduce, HDFS, HBase, Mongo-DB, CentOS 6.4, Unix, REST web Services, ANT 1.6, Elastic Search, Hive, Pig, Oozie, Java, JSON, Eclipse, Qlik view, Qlik Sense, Oracle Database, Jenkins, Maven, Sqoop.
Confidential
Java/J2EE Developer
Responsibilities:
- Involved in Requirements Analysis and design an Object-oriented domain model.
- Involvement in the detailed Documentation, written functional specifications of the module.
- Involved in development of Application with Java and J2EE technologies.
- Develop and maintain elaborate services-based architecture utilizing open source technologies like Hibernate, ORM and Spring Framework.
- Designed and documented REST/HTTP APIs, including JSON data formats and API versioning strategy.
- Developed server-side services using Java multithreading, Struts MVC, Java, EJB, Spring, Web Services (SOAP, WSDL, AXIS).
- Used Micro Services as communicating medium for different APIs, processed large number of small Processes.
- Involvement in creating and configuring of build files using Ant.
- Development of Controller Servlet a Framework component for Presentation.
- Investigated MVC framework technologies including JSF based (ICEfaces,RichFaces) and to implement the MVC architecture of the product.
- Developed application using JSF, myFaces, Spring, and JDO technologies which communicated with Mainframe software.
- Designing, Development and Implementation of JSPs in Presentation layer for Submission, Application, reference implementation.
- Developed JSP pages using Jdeveloper9.0.5, such as HTML, Bean Tags, Logic Tags and Template Tags.
- Practiced and evangelized agile development approaches. Wrote ANT scripts and assisted with build and configuration management processes.
- Development of JavaScript for client end data entry validations and Front-End Validation.
- Deployed Web, presentation and business components on Apache Tomcat Application Server.
- Generating schema difference reports for database using toad.
- Developed PL/SQL procedures for different use case scenarios
- Built the report module on reports based from Crystal reports.
- Involvement in post-production support, Testing and used JUNIT for unit testing of the module.
Environment: Java/J2EE, JSP, XML, Spring Framework, Hibernate, Eclipse (IDE), Micro Services, Java Script, Struts, Tiles, Ant, SQL, PL/SQL, Oracle, Windows, UNIX, Soap, Jasper reports.
Confidential
Junior Java Developer
Responsibilities:
- Actively involved from fresh start of the project, requirement gathering to quality assurance testing.
- Coded and Developed Multi-tier architecture in Java, J2EE, Servlets.
- Conducted analysis, requirements study and design according to various design patterns and developed rendering to the use cases, taking ownership of the features.
- Used various design patterns such as Command, Abstract Factory, Factory, and Singleton to improve the system performance. Analyzing the critical coding defects and developing solutions.
- Developed configurable front-end using Struts technology. Also involved in component-based development of certain features which were reusable across modules.
- Designed, developed and maintained the data layer using the ORM framework called Hibernate.
- Used Hibernate framework for Persistence layer, involved in writing Stored Procedures for data retrieval and data storage and updates in Oracle database using Hibernate.
- Developed batch jobs which will run on specified time to implement certain logic in java platform.
- Developing & deploying Archive files (EAR, WAR, JAR) using ANT build tool.
- Used Software development best practices for Object Oriented Design and methodologies throughout Object oriented development cycle
- Responsible for developing SQL Queries required for the JDBC.
- Designed the database and worked on DB2 and executed DDLS and DMLS.
- Active participation in architecture framework design and coding and test plan development.
- Strictly followed Water Fall development methodologies for implementing projects.
- Thoroughly documented the detailed process flow with UML diagrams and flow charts for distribution across various teams.
- Involved in developing training presentations for developers (off shore support), QA, Production support.
- Presented the process logical and physical flow to various teams using PowerPoint and Visio diagrams.
Environment: Java JDK (1.5), Java J2EE, Informatica, Oracle 11g (TOAD and SQL developer) Servlets, Jboss application Server, Water Fall, JSPs, EJBs, DB2, RAD, XML, Web Server, JUNIT, Hibernate, MS ACCESS, Microsoft Excel.
