Sr. Java Bigdata Developer Resume
Reston, VA
SUMMARY:
- Over 10+ years of experience in Information Technology which includes experience in Big data, HADOOP Ecosystem, Core Java/J2EEand strong in Design, Software processes, Requirement gathering, Analysis and development of software applications
- Excellent Hands on Experience in developing Hadoop Architecture in Windows and Linux platforms.
- Experience in building bigdata solutions using Lambda Architecture using Cloudera distribution of Hadoop, Twitter Storm, Trident, MapReduce, Cascading, HIVE, PIG and Sqoop.
- Strong development experience in Java/JDK 7, JEE6, Maven, Jenkins, Jersey, Servlets, JSP, Struts, Spring, Hibernate, JDBC, Java Beans, JMS, JNDI, XML, XML Schema, Web Services, SOAP, JUnit, ANT, Log4j, HTML, JavaScript, Node JS, React JS, Angular Js.
- Expertise in various components of Hadoop Ecosystem - Map Reduce, Hive, Pig, Sqoop, Impala, Flume, Oozie, HBase, MongoDb, Cassandra, Scala, Spark, Kafka, YARN.
- Experienced in J2EE Design Patterns such as MVC, Business Delegate, Service Locator, Singleton, Transfer Object, Singleton, Session Façade, and Data Access Object.
- Good Knowledge in Amazon Web Service (AWS) concepts like EMR and EC2 web services which provides fast and efficient processing of Teradata Big Data Analytics.
- Experience working with Amazon AWS cloud which includes services like (EC2, S3, RDS and EBS), Elastic Beanstalk, Cloud Watch.
- Hands on experience in installing, configuring, and using Hadoop ecosystem components like HDFS, Hive, Spark, Scala, Spark-SQL, MapReduce, Pig, Sqoop, Flume, HBase, Zookeeper, and Oozie.
- Excellent working experience on Big Data Integration and Analytics based on Hadoop, SOLR, Spark, Kafka, Storm and web Methods technologies.
- Experienced in designing and developing applications in Spark using Scala to compare the performance of Spark with Hive and SQL/Oracle.
- Experienced in collection of Log Data and JSON data into HDFS using Flume and processed the data using Hive/Pig.
- Hands on experience working on NoSQL databases including Hbase, MongoDB, Cassandra and its integration with Hadoop cluster.
- Strong Knowledge and experience on implementing Big Data in Amazon Elastic MapReduce (Amazon EMR) for processing, managing Hadoop framework dynamically scalable Amazon EC2 instances.
- Hands on experience in writing Ad-hoc Queries for moving data from HDFS to HIVE and analyzing the data using HIVEQL.
- Good knowledge in RDBMS concepts (Oracle 12c, 11g, MS SQL Server 2016) and strong SQL, PL/SQL query writing skills (by using TOAD & SQL Developer tools), Stored Procedures and Triggers.
- Expertise in Amazon Web Services including Elastic Cloud Compute (EC2) and Dynamo DB and expertise in Automating deployment of large Cassandra Clusters on EC2 using EC2 APIs
- Experienced in developing, deploying enterprise applications on IBM Web Sphere, BEA WebLogic, Oracle Application Server, JBoss, Tomcat, and Jetty.
- Experienced in development and utilization of Apache SOLR with Data Computations and Transformation for use by Down Stream Online Applications.
- Experienced in importing and exporting data using Sqoop from HDFS (Hive & HBase) to Relational Database Systems (Oracle &Teradata) and vice-versa.
- Expertise in using IDE like WebSphere (WSAD), Eclipse, NetBeans, MyEclipse, WebLogic Workshop and experienced in developing and designing Web Services (SOAP and Restful Web services).
- Highly Proficient in writing complex SQL Queries, stored procedures, triggers and very well experienced in PL/SQL or T-SQL.
- Proficient in developing Web based user interfaces using HTML5, CSS3, JavaScript, jQuery, AJAX, XML, JSON, jQuery UI, Bootstrap, AngularJS, Node JS, and Ext JS.
- Expertise in various Java/J2EE technologies like JSP, Servlets, Hibernate, Struts, spring.
- Good knowledge with web-based UI development using jQuery UI, jQuery, ExtJS, CSS3, HTML, HTML5, XHTML and JavaScript.
TECHNICAL SKILLS:
Big Data:: Hadoop, Storm, Hbase, Hive, Flume, Cassandra, Kafka, Storm, Sqoop, Oozie, PIG, Spark, MapReduce, ZooKeeper, Yarn, MongoDB, Cassandra, Cloudera.
Java/J2EE Techs: Spring, Hibernate, Struts, JSP, HTML, CSS, AJAX, JavaScript, JASON, Angular.js, JQuery, Servlets, EJB, Web Services, SOAP, Restful, XML, DHTML.
Operating Systems:: UNIX, Mac, Linux, Windows 2000 / NT / XP / Vista, Android
Programming Languages:: Java (JDK 5/JDK 6&7), R, HTML, SQL, PL/SQL
Frameworks:: Hibernate 2.x/3.x, Spring 2.x/3.x,Struts 1.x/2.x and JPA
Web Services:: WSDL, SOAP, Apache CXF/XFire, Apache Axis, REST, Jersey
Databases:: Oracle 10g/11g/12c, Microsoft SQL Server, DB2 & MySQL 4.x/5.x
Middleware Technologies:: Web sphere Message Queue, Web sphere Message Broker, XML gateway, JMS
Web Technologies:: J2EE, Soap & REST Web Services, JSP, Servlets, EJB, JavaScript, Struts, spring, web works, HTML, XML, JMS, JSF and Ajax.
Testing Frameworks:: Mockito, PowerMock, EasyMock
Web/Application Servers:: IBM Web sphere Application server, JBoss, Apache Tomcat
Others:: Software Borland Star team, Clear case, Junit, ANT, Maven, Android Platform, Microsoft Office, SQL Developer, DB2 control center, Microsoft Visio, Hudson, Subversion, GIT, Nexus, Arti factory
Development Strategies:: Agile, Lean Agile, Pair Programming, Water-Fall and Test Driven Development
PROFESSIONAL EXPERIENCE:
Confidential - Reston, VA
Sr. Java BigData Developer
Responsibilities:
- Gathered the business requirements from the Business Partners and Subject Matter Experts and involved in installation and configuration of Hadoop Ecosystem components with Hadoop Admin.
- Working on architected solutions that process massive amounts of data on corporate and AWS cloud based servers.
- Supported MapReduce Programs those are running on the cluster and also wrote MapReduce jobs using Java API.
- Configure a number of node (Amazon EC2 spot Instance) Hadoop cluster to transfer the data from Amazon S3 to HDFS and HDFS to AmazonS3 and also to direct input and output to the Hadoop MapReduce framework.
- Involved in HDFS maintenance and loading of structured and unstructured data and imported data from mainframe dataset to HDFS using Sqoop.
- Handled importing of data from various data sources (i.e. Oracle, DB2, HBase, Cassandra, and MongoDB) to Hadoop, performed transformations using Hive, MapReduce.
- Developed prototype Spark applications using Spark-Core, Spark SQL, Data Frame API and developed several custom User defined functions in Hive & Pig using Java & python
- Developed simple and complex MapReduce programs in Java for Data Analysis on different data formats.
- Importing the data into Spark from Kafka Consumer group using Spark Streaming APIs.
- Wrote Hive queries for data analysis to meet the business requirements. Implemented Kafka Custom encoders for custom input format to load data into Kafka Partitions. Real time streaming the data using Spark with Kafka for faster processing.
- Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scala and worked on Kafka, Kafka-Mirroring to ensure that the data is replicated without any loss.
- Written python scripts for internal testing which pushes the data reading form a file into Kafka queue which in turn is consumed by the Storm application.
- Installed and configured Flume, Hive, Pig, Sqoop and Oozie on the Hadoop cluster.
- Participated in building CDH4 test cluster for implementing Kerberos authentication. Upgraded the Hadoop Cluster from CDH4 to CDH5 and setup High availability Cluster to Integrate the HIVE with existing applications
- Implemented AWS EC2, Key Pairs, Security Groups, Auto Scaling, ELB, SQS, and SNS using AWS API and exposed as the Restful Web services and implemented Reporting, Notification services using AWS API.
- Load and transform large sets of structured, semi structured and unstructured data using Hadoop/Big Data concepts.
- Involved in migrating Hive queries into Spark transformations using Data frames, Spark SQL, SQL Context, and Scala.
- Worked with various HDFS file formats like Avro1.7.6, Sequence File, Jsonandvarious compression formats like Snappy, bzip2.
- Implemented Data Interface to get information of customers using RestAPI and Pre-Process data using MapReduce 2.0 and store into HDFS (Hortonworks)
- Responsible for migrating tables from traditional RDBMS into Hive tables using Sqoop and later generate required visualizations and dashboards using Tableau and generate final reporting data using Tableau for testing by connecting to the corresponding Hive tables using Hive ODBC connector.
- Worked on documentation of all Extract, Transform and Load, designed, developed, validated and deploy the Talend ETL processes for Data ware house team using PIG, HIVE.
- Implemented Storm builder topologies to perform cleansing operations before moving data into Cassandra and prototype done with HDP Kafka and Storm for click stream application.
- Developed the Pig 0.15.0UDF's to pre-process the data for analysis and Migrated ETL operations into Hadoop system using Pig Latin scripts and Python Scripts3.5.1.
- Used AWS (Amazon Web services) compute servers extensively and create Snapshots of EBS Volumes. Monitor AWS EC2 Instances using Cloud Watch.
- Implemented Daily Oozie jobs that automate parallel tasks of loading the data into HDFS and pre-processing with Pig using Oozie co-coordinator jobs.
- Involved in importing and exporting data between HDFS and Relational Database Systems like Oracle, MySQL and SQL Server using Sqoop.
- Updated maps, sessions and workflows as a part of ETL change and also modified existing ETL Code and document the changes.
Environment: Hadoop, Java, J2EE, MapReduce, Python, HDFS, Hbase, Hive, Pig, Linux, XML, Eclipse, Kafka, Storm, Spark, Cloudera, CDH4/5 Distribution, DB2, Scala, SQL Server, Oracle 12c, MySQL, Talend, MOngoDB, Cassandra, AWS, Tableau, Oozie, Restful, SOAP and JavaScript.
Confidential - Mentor, OH
Sr. Java Big Data/Hadoop Developer
Responsibilities:
- Worked on analyzing Hadoop cluster using different big data analytic tools including Flume, Pig, Hive, HBase, Oozie, Zookeeper, Sqoop, Spark and Kafka.
- Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data and used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
- As a Big Data Developer implemented solutions for ingesting data from various sources and processing the Data-at-Rest utilizing Big Data technologies such as Hadoop, MapReduce Frameworks, MongoDB, Hive, Oozie, Flume, Sqoop and Talend etc.
- Developed a job server (REST API, spring boot, ORACLE DB) and job shell for job submission, job profile storage, job data (HDFS) query/monitoring.
- Explored with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark -SQL, Data Frame, Pair RDD's, Spark YARN.
- Deployed application to AWS and monitored the load balancing of different EC2 instances
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and Extracted the data from SQL into HDFS using Sqoop.
- Created Elastic Map Reduce (EMR) clusters and Configured the Data pipeline with EMR clusters for scheduling the task runner and provisioning of Ec2 Instances on both Windows and Linux.
- Developed analytical components using Scala, Spark, Apache Mesos and Spark Stream.
- Worked on AWS Relational Database Services, AWS Security Groups and their rule and implemented Reporting, Notification services using AWS API.
- Installed Hadoop, Map Reduce, and HDFS and developed multiple MapReduce jobs in PIG and Hive for data cleaning and pre-processing.
- Worked on Big Data Integration &Analytics based on Hadoop, SOLR, Spark, Kafka, Storm and web Methods.
- Developed Kafka producer and consumers, Spark and Hadoop MapReduce jobs and imported the data from different sources like HDFS/Hbase into Spark RDD.
- Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
- Strongly recommended to bring in Elastic Search and was responsible for installing, configuring and administration.
- Implemented AWS EC2, Key Pairs, Security Groups, Auto Scaling, ELB, SQS, and SNS using AWS API and exposed as the Restful Web services.
- Involved in converting MapReduce programs into Spark transformations using Spark RDD's on Scala.
- Configured deployed and maintained multi-node Dev and Test Kafka Clusters and developed Spark scripts by using Scala Shell commands as per the requirement.
- Performed transformations, cleaning and filtering on imported data using Hive, Map Reduce, and loaded final data into HDFS.
- Implemented using SCALA and SQL for faster testing and processing of data. Real time streaming the data using with KAFKA.
- Design & implement ETL process using Talend to load data from Worked extensively with Sqoop for importing and exporting the data from HDFS to Relational Database systems/mainframe and vice-versa. Loading data into HDFS.
- Load the data into Spark RDD and do in memory data Computation to generate the Output response.
- Worked on major components in Hadoop Ecosystem including Hive, PIG, HBase, HBase-Hive Integration, Scala, Sqoop and Flume.
- Developed Hive Scripts, Pig scripts, UNIX Shell scripts, programming for all ETL loading processes and converting the files into parquet in the Hadoop File System.
- Worked with Oozie and Zookeeper to manage job workflow and job coordination in the cluster and developed and written Apache PIG scripts and HIVE scripts to process the HDFS data.
- Used Hive to find correlations between customer's browser logs in different sites and analyzed them to build risk profile for such sites.
- Utilized Agile Scrum Methodology to help manage and organize a team of 4 developers with regular code review sessions.
- Created and maintained Technical documentation for launching Hadoop Clusters and for executing Hive queries and Pig Scripts.
Environment: Hadoop, J2EE, JavaScript, HDFS, Spark, MapReduce, Pig, Hive, Sqoop, Kafka, HBase, Oozie, Flume, Scala, Python, Java, SQL Scripting and Talend, Linux Shell Scripting, Cassandra, Zookeeper, HBase, MongoDB, Cloudera, Cloudera Manager, EC2, EMR, S3, Oracle, MySQL.
Confidential - Johns Creek, GA
Sr. Java Big Data/Hadoop Developer
Responsibilities:
- Worked on analyzing Hadoop cluster using different big data analytic tools including Kafka, Pig, Hive and MapReduce.
- Proactively monitored systems and services, architecture design and implementation of Hadoop deployment, configuration management, backup, and disaster recovery systems and procedures
- Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scale and handled in Importing and exporting data into HDFS and Hive using SQOOP and Kafka.
- Installed and configured Hadoop, MapReduce, HDFS (Hadoop Distributed File System), developed multiple MapReduce jobs in java for data cleaning and processing.
- Designed and configured Flume servers to collect data from the network proxy servers and store to HDFS and HBASE.
- Worked on implementing Spark using Scala and Spark SQL for faster analyzing and processing of data.
- Used JAVA, J2EE application development skills with Object Oriented Analysis and extensively involved throughout Software Development Life Cycle (SDLC) and utilized Java and MySQL from day to day to debug and fix issues with client processes
- Involved in launching and Setup of HADOOP/ HBASE Cluster which includes configuring different components of HADOOP and HBASE Cluster.
- Hands-on experience of Web logic Application Server, Web Sphere Application Server, Web Sphere Portal Server, and J2EE application deployment technology
- Involved in HDFS maintenance and WEBUI it through Hadoop-Java API and involved in creating Hive tables, loading the data and writing hive queries, which will run internally in a map reduce.
- Applied MapReduce framework jobs in java for data processing by installing and configuring Hadoop, HDFS.
- Involved in developing Pig Scripts for change data capture and delta record processing between newly arrived data and already existing data in HDFS.
- Extensively used Sqoop to get data from RDBMS sources like Teradata and Netezza and worked on importing the unstructured data into the HDFS using Flume.
- Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Worked on Designing and Developing ETL Workflows using Java for processing data in HDFS/Hbase using Oozie and involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs
- Involved in developing Shell scripts to easy execution of all other scripts (Pig, Hive, and MapReduce) and move the data files within and outside of HDFS.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
- Worked with NoSQL databases like Hbase in creating tables to load large sets of semi structured data and worked on loading data from UNIX file system to HDFS
- Generated JavaAPIs for retrieval and analysis on No-SQL database such as HBase.
- Created ETL jobs to generate and distribute reports from MySQL database using Pentaho Data Integration.
Environment: Hadoop, Java/J2EE, HDFS, MapReduce, Hive Sqoop, Pig, Hbase, Apache Spark, Oozie Scheduler, Java, UNIX Shell Scripts, Kafka, Git, Maven, PLSQL, MongoDB, HBase, Cassandra, Python, Scala, Teradata, Netezza, Oracle.
Confidential - Baltimore, MD
Sr. Java/Hadoop Developer
Responsibilities:
- Installed/Configured/Maintained Apache Hadoop clusters for application development and Hadoop tools like Hive, Pig, HBase, Flume, Oozie Zookeeper and Sqoop.
- Involved in writing Client side Scripts using Java Scripts and Server Side scripts using Java Beans and used servlets for handling the business.
- Developed applications in Hadoop Big Data technologies- Pig, Hive, Map-Reduce, Hbase and Oozie.
- Implemented REST Web services using JAX-RS (Jersey) by Spring Framework and consumption of data using AJAX.
- Developed Scala programs with Spark for data in Hadoop ecosystem.
- Extensively involved in Installation and configuration of Cloudera distribution Hadoop 2, 3, NameNode, Secondary NameNode, JobTracker, TaskTrackers and DataNodes.
- Develop test cases using Java test frameworks such as Junit, Power mock, Easy mock and Spring Test for entire application modules.
- Developed another user based Web services (SOAP) through WSDL using WebLogic application server and JAXB as binding framework to interact with other components.
- Implemented design patterns such as Data Access Object, Data Transfer Object, Business Delegate, Session Facade, Service Locator and Singleton.
- Developed code for handling bean references in spring framework using Dependency Injection (DI) Inversion of Control (IOC) using annotations.
- Managed and reviewed Hadoop Logfiles as a part of administration for troubleshooting purposes. Communicate and escalate issues appropriately and developed MapReduce jobs using apache commons components.
- Used Service Oriented Architecture (SOA) based SOAP and REST Web Services (JAX-RS) for integration with other systems.
- Developed components of web services (JAX-WS, REST, JAX-RPC) end-to-end, using different JAX-WS standards with clear understanding on WSDL (type, message, port Type, bindings, and service).
- Collected and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis
- Involved in designing and developing the application using JSTL, JSP, Java script, AJAX, HTML, CSS and collection.
- Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
- Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and translate to MapReduce jobs.
- Implemented Spring ORM with Hibernate taking advantage of Java features like annotation metadata, auto wiring, and generic collections that is used to implement the DAO layer with Hibernate Entity Manager’s Session Factory, HQL, and SQL.
- Developed UDFs in Java as and when necessary to use in PIG and HIVE queries.
- Developed Java Web Applications using JSP and Servlets, Struts, Hibernate, spring, Rest Web Services, SOAP.
- Worked over the entire Software Development Life Cycle (SDLC) as a part of a team as well as independently.
- Implemented MAVEN builds tools for packaging and deployment on a target server.
- Written SQL queries to query the database and providing data extracts to users as per request.
Environment: Java 1.5, JSP, Servlet, Spring, Hibernate 3.0, Struts framework, Hadoop, Map Reduce, HDFS, HBase, Hive, Pig, Sqoop, Flume, Kafka, Spark, Scala, ETL, Cloudera CDH Apache Hadoop, HTML, XML, Log 4j, Eclipse, Unix, Windows XP
Confidential - Brentwood, TN
Java/J2EE Developer
Responsibilities:
- Designed and implemented the strategic modules like Underwriting, Requirements, Create Case, User Management, Team Management and Material Data Changes and developed web pages using Struts, JSP, Servlets, HTML and JavaScript.
- Provide support in all phases of Software development life cycle (SDLC), quality management systems and project life cycle processes. Utilizing Database Such as MYSQL, Following HTTP and WSDL Standards to Design the REST/ SOAP Based Web API’S using XML, JSON, HTML, and DOM Technologies.
- Involved in Installation and Configuration of Tomcat, Spring Source Tool Suit, Eclipse, unit testing.
- Back end server side coding and development using Java data structure as a Collections including Set, List, Map, Exception Handling, Spring with dependency injection, Struts Framework, Hibernate, Servlets, Action, Action Forms &Java beans, etc.
- Implemented Singleton, Service Locator design patterns in MVC framework and developed command, delegate, model action script classes to interact with the backend.
- Involved in Migrating existing distributed JSP framework to Struts Framework, designed and involved in research of Struts MVC framework
- Responsible for Web UI development in JavaScript using JQuery, AngularJs, and AJAX and developed Ajax framework on service layer for module as benchmark
- Implemented Model View Controller (MVC) Architecture and coded Java Beans (as the model) and implemented Service and DAO layers in between Struts and Hibernate.
- Designed Graphical User Interface (GUI) applications using HTML, JSP, JavaScript (JQuery), CSS and AJAX.
- Applied MVC pattern of Ajax framework which involves creating Controllers for implementing Classes.
- Developed Spring REST Web services for opening, closing the locker door Web service operations.
- Responsible to enhance the UI using HTML, Java Script, XML, JSP, CSS as per the requirements and providing the client side using JQuery validations.
- Involved in write application level code to interact with APIs, Web Services using AJAX, JSON and XML.
- Wrote lots of JSP's for maintains and enhancements of the application. Worked on Front End using Servlets, JSP and also backend using Hibernate.
- Used Hibernate, object/relational-mapping (ORM) solution, and technique of mapping data representation from MVC model to Microsoft Relational data model with a T-SQL based schema.
- Front end development utilizing HTML5, CSS3, and JavaScript leveraging the Bootstrap framework and a Java backend
- Used JAXB for converting Java Object into a XML file and for converting XML content into a JavaObject.
- Web services were built using Spring and CXF operating within Mule ESB; offering both REST and SOAP interfaces.
- Involved in Maven Script to create JAR, WAR, EAR & dependency JARS and used Maven as the build tool for the application and used JIRA for bug/task tracking and time tracking.
- Used agile methodology for development of the application.
Environment: Java J2EE, JSP, JavaScript, Ajax, Swing, Spring 3.2, Eclipse 4.2, Hibernate 4.1, XML, Tomcat, Oracle 10g, JUnit, JMS, Log4j, Maven, Agile, Git, JDBC, Web service, XML, SOAP, JAX-WS, Unix MongoDB, AngularJS and Soap UI.
