Sr.java-bigdata Developer Resume
Milwaukee, WI
PROFESSIONAL SUMMARY:
- 8+ years of professional experience this includes Analysis , Design , Development , Integration , Deployment and Maintenance of quality softwareapplications using Java / J2EE Technologies and Hadoop technologies.
- Experienced in installing, configuring, testing Hadoop ecosystem components on Linux / UNIX including Hadoop Administration (like Hive, pig, Sqoop etc.)
- Expertise in Java, Hadoop MapReduce , Pig , Hive , Oozie , Sqoop , Flume , Zookeeper , Impala and NoSQL Database.
- Excellent experience on Hadoop ecosystem, In - depth understanding of Map Reduce and the Hadoop Infrastructure.
- Excellent experience in Amazon , Cloudera and Hortonworks Hadoop distribution and maintaining and optimized AWS infrastructure ( EMREC2 , S3 , EBS )
- Expertise in developing Spark code using Scala and Spark - SQL /Streaming for faster testing and processing of data.
- Excellent knowledge on Hadoop Architecture and ecosystems such as HDFS , JobTracker , TaskTracker , NameNode , Data Node , YARN and MapReduce programming paradigm.
- Experienced working with Hadoop Big Data technologies (hdfs and Mapreduce programs), Hadoop ecosystems
- ( Hbase , Hive , pig ) and NoSQL database MongoDB
- Experienced on usage of NoSQL database column-oriented Cassandra .
- Extensive experience in working with semi / unstructureddata by implementing complex map reduce programs using design patterns.
- Experienced on major components in Hadoop Ecosystem including Hive, Sqoop, Flume &knowledge of MapReduce / HDFSFramework .
- Experienced in working with MapReduce Design patterns to solve complex MapReduce programs.
- Excellent Knowledge in Talend Big data integration for business demands to work towards Hadoop and NoSQL
- Hands-on programming experience in various technologies like JAVA , J2EE , HTML , XML
- Excellent Working Knowledge on Sqoop and Flume for Data Processing
- Expertise in loading the data from the different data sources like ( Teradata and DB2 ) into HDFS using Sqoop and load into partitioned Hive tables
- Experienced on Hadoop cluster maintenance including data and metadata backups, file system checks, commissioning and decommissioning nodes and upgrades.
- Extensive experience writing custom Map Reduce programs for data processing and UDFs for both Hive and Pig in Java .
- Strong experience in analysing large amounts of data sets writing Pigscripts and Hivequeries .
- Extensive experienced in working with structured data using HiveQL , join operations, writing custom UDF's and experienced in optimizing HiveQueries .
- Experienced in importing and exporting data using Sqoop from HDFS to Relational Database.
- Expertise in job workflow scheduling and monitoringtools like Oozie .
- Experience in ApacheFlume for collecting , aggregating and moving huge chunks of data from various sources such as webserver , telnet sources etc.
- Extensively designed and executed SQL queries to ensure data integrity and consistency at the backend.
- Strong experience in architecting batch style large scale distributedcomputingapplications using tools like Flume , Mapreduce , hive etc.
- Experience using various Hadoop Distributions (Cloudera, Hortonworks, MapR etc) to fully implement and leverage new Hadoop features
- Worked on custom PigLoaders and Storageclasses to work with a variety of data formats such as JSON , CompressedCSV .
- Experience in working with different scriptingtechnologies like Python , UNIXshellscripts .
- Strong experience in working with UNIX / LINUX environments, writingshellscripts .
- Excellent knowledge and working experience in Agile & Waterfall methodologies.
- Designing and developing data warehouse and RedshiftBI Based solutions.
- Involved in Designing AmazonRedshift DB clusters, schema and tables.
- Involved in writing complex SQL’s using windowing functions to extract data from redshift without stored Procedure.
- Applied MachineLearning and performed statisticalanalysis on the data.
- Scraped and analyzed data using MachineLearning algorithms in Python and SQL.
- Expertise in Web pages development using JSP , HTML , JavaScript , jQuery and Ajax .
- Experience in writing database objects like StoredProcedures , Functions , Triggers , PL / SQL packages and Cursors for Oracle , SQLServer , and MySQL & Sybase databases.
- Great team player and quick learner with effectivecommunication , motivation , and organizational skills combined with attention to details and business improvements.
TECHNICALSKILLS :
- Hadoop/Big Data, HDFS, Map Reduce, Hive, Pig, Sqoop, Flume, Oozie, Scala, spark, storm, Kafka, Rabbit MQ, Active MQ, Zoo Keeper.
- HBase, Cassandra, CouchDB, MongoDB.
- Cloudera, HortonWorks, MapR.
- Teradata, MS SQL Server, Oracle, Informix, Sybase, Informatica, Datastage.
- JAVA, J2EE, Spring, Hibernate EJB, Webservices (JAX - RPC, JAXP, JAXM), JMS, JNDI, Servlets, JSP, Jakarta Struts, Python.
- BEA Web Logic, IBM Websphere, JBoss, Tomcat.
- UML, OOAD.
- HTML, AJAX, CSS, XHTML, XML, XSL, XSLT, WSDL, JSON, SOAP, REST, GRAILS
- CVS, SVN, SharePoint, Clear Case, Clear Quest, Win CVS, Junit, MRUnit, Ant, Maven, Log4j, FrontPage
- Eclipse, NetBeans.
- Linux, UNIX, Windows
WORK EXPERIENCE:
Sr.Java-BigData Developer
Confidential, Milwaukee, WI
Responsibilities:
- Involved in full life cycle of the project from Design , Analysis , logical and physical architecture modelling , development , Implementation , testing
- Creating Nifi Processor to move Cassandra data into Mysql Core Ontology.
- Developed a POC Environment Using NIFI to Orchestrate the workflow from Kafka to Mysql.
- Reverse Engineering Informatica into Java using Custom Nifi Processors.
- Using Kafka for Durable Event and Batch Communication.
- Developing Orchestration work to hit multiple endpoints for rest services using Java, Vert.x and RxJava.
- Consuming RESTAPI’S to see RequestAndResponse.
- Replicating Vendor API Model to Own NM Legacy Model.
- Creating data model Objects and converting the request and response from the Vendor API Model and Hitting it through theVerticle to consume theRestSevice.
- Running theRESTAPI Model through CICD Pipeline Using Kubernetes Cluster.
- Converting XML Mappings in informatica to Java.
- Understand the requirement and write test cases.
- Integrate tests with SonarQube and monitor the code coverage.
- Validated RestFulAPI Services.
- Designed and Documented RESTAPI’S Including JSON Data Formats and API Versioning Strategy.
Environment : Hadoop, HDFS, Pig, Sqoop, Mysql, Maven, Gradle, Shell Scripting, CDH, Cassandra, Cloudera, AWS (S3, EMR), SQL, Kubernetes, Spark, RDBMS, Java, HTML, Java, Nifi, JavaScript, WebServices, CICD, Redshift, Kafka, NiFi, SonarQube, Microservices.
Sr. Big Data/Hadoop Developer.
Confidential, Norwalk,CT
Responsibilities:
- Involved in full life cycle of the project from Design , Analysis , logical and physical architecture modelling , development , Implementation , testing .
- Scripts were written for distribution of query for performance test jobs in AmazonDatalake .
- Created Hive Tables, loaded transactional data from Teradata using Sqoop and Worked with highly unstructured and semi structured data of 2 Petabytes in size .
- Developed MapReduce ( YARN ) jobs for cleaning , accessing and validating the data.
- Created and worked Sqoopjobs with incremental load to populate Hive External tables.
- Developed optimal strategies for distributing the weblogdata over the cluster importing and exporting the stored web logdata into HDFS and Hive using Sqoop .
- ApacheHadoop installation & configuration of multiple nodes on AWSEC2 system
- Developed PigLatinscripts for replacing the existing legacyprocess to the Hadoop and the data is fed to AWSS3 .
- Responsible for building scalabledistributeddata solutions using HadoopCloudera .
- Designed and developed automation test scripts using Python
- Integrated Apache Storm with Kafka to perform webanalytics and to perform clickstreamdata from Kafka to HDFS .
- Writing Pig - scripts to transform raw data from several data sources into forming baselinedata .
- Analysed the SQL scripts and designed the solution to implement using Pyspark
- Implemented Hive Generic UDF's to in corporate businesslogic into HiveQueries .
- Responsible for developing data pipeline with AmazonAWS to extract the data from weblogs and store in HDFS .
- Uploaded streaming data from Kafka to HDFS , Cassandra and Hive by integrating with storm .
- Analysed the web log data using the HiveQL to extract number of unique visitors per day, page views, visit duration, most visited page on website.
- Supporting data analysis projects by using Elastic MapReduce on the AmazonWebServices ( AWS ) cloud performed Export and import of data into s3 .
- Worked on MongoDB by using CRUD ( Create , Read , Update and Delete ), Indexing , Replication and Sharding features.
- Involved in designing the rowkey in Hbase to store Text and JSON as key values in Hbase table and designed row key in such a way to get/scan it in a sorted order. • Integrated Oozie with the rest of the Hadoop stack supporting several types of Hadoop jobs out of the box (such as Map - Reduce , Pig , Hive , and Sqoop ) as well as system specific jobs (such as Javaprograms and shellscripts ). • Worked on custom talend jobs to ingest , enrich and distribute data in Cloudera Hadoop ecosystem . • Creating Hive tables and working on them using HiveQL . • Designed and Implemented Partitioning (Static, Dynamic) Buckets in HIVE . • Developed multiple POCs using PySpark and deployed on the YARN cluster, compared the performance of Spark, with Hive and SQL and Involved in End-to-End implementation of ETL logic. • Developed syllabus/Curriculum data pipelines from Syllabus/Curriculum WebServices to Cassandra and Hive tables.
- Ingested Streaming data with Apache NIFI into Kafka .
- Worked with NIFI for managing flow of data from sources through automateddata flow.
- Developed and deployed Apache NIFI flows across various environments, optimized Nifi data flows.
- Involved in Stream processing using spark and flink on the data ingesting using kafka and build analytical models out of that for downstream applications. • Worked on Cluster co-ordination services through Zookeeper . • Monitored workload, job performance and capacity planning using ClouderaManager . • Exported the analyzed data to the RDBMS using Sqoop for to generate reports for the BI team.
- Migrating existing application into REST based microservices to provide all the CRUD capabilities using Spring Boot.
- Creating the cube in talend to create different types of aggregation in the data and also to visualize them.
- Involved in Agilemethodologies , daily scrummeetings , spring planning.
Environment : Hadoop, HDFS, Map Reduce, Hive, Pig, Sqoop, Oozie, Maven, Python, Shell Scripting, CDH, MongoDB, HBase, Cloudera, AWS (S3, EMR), SQL, Python, Scala, Spark, RDBMS, Java, HTML, Pyspark, Nifi, JavaScript, WebServices,flink, Redshift, Kafka, Strom, Talend, Microservices.
Hadoop Developer
Confidential, Cincinnati, OH
Responsibilities:
- Responsible for installation and configuration of Hive , Pig , Hbase and Sqoop on the Hadoopcluster and created hive tables to store the processed results in a tabularformat .
- Configured Spark Streaming to receive real time data from the ApacheKafka and store the stream data to HDFS using Scala .
- Developed the Sqoopscripts in order to make the interaction between Hive and verticaDatabase .
- Processed data into HDFS by developing solutions and analyzed the data using MapReduce , PIG , and Hive to produce summary results from Hadoop to downstreamsystems .
- Build servers using AWS : Importing volumes, launching EC2 , creating security groups , auto - scaling, load balancers, Route53 , SES and SNS in the defined virtualprivate connection.
- Written Map Reduce code to process and parsing the data from various sources and storing parsed data into HBase and Hive using HBase - HiveIntegration .
- Streamed AWS log group into Lambda function to create servicenow incident.
- Involved in loading and transforming large sets of Structured , Semi - Structured and Unstructured dataand analyzed them by running Hivequeries and Pigscripts .
- Created Managed tables and External tables in Hive and loaded data from HDFS .
- Developed Spark code by using Scala and Spark - SQL for faster processing and testing and performed complex HiveQL queries on Hive tables.
- Scheduled several Time-based Oozie workflow by developing Python scripts.
- Developed Pig Latin scripts using operators such as LOAD , STORE , DUMP , FILTER , DISTINCT , FOREACH , GENERATE , GROUP , COGROUP , ORDER , LIMIT , UNION , SPLIT to extract data from data files to load into HDFS .
- Exporting the data using Sqoop to RDBMS servers and processed that data for ETL operations.
- Worked on S3buckets on AWS to store Cloud Formation Templates and worked on AWS to create EC2 instances.
- Designing ETL Data Pipeline flow to ingest the data from RDBMS source to Hadoop using shell script, sqoop , package and MySQL .
- End - to - end architecture and implementation of client-server systems using Scala , Java , JavaScript and related, Linux
- Used Hadoop sPig, Hive and Map Reduce for analysing the Healthinsurance data to help by extracting data sets for meaningful information such as medicines, diseases, symptoms, opinions, geographic region detail etc.
- Optimized the Hive tables using optimization techniques like partitions and bucketing to provide better.
- Used Oozieworkflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Javamap - reduceHive , Pig , and Sqoop .
- Implementing Hadoop with the AWSEC2 system using a few instances in gathering and analysing data logfiles .
- Involved in Spark and Spark Streaming creating RDD's , applying operations - Transformation and Actions .
- Created partitioned tables and loaded data using both staticpartition and dynamicpartition method.
- Developed custom ApacheSpark programs in Scala to analyse and transform unstructured data.
- Handled importing of data from various data sources, performed transformations using Hive , MapReduce , loaded data into HDFS and Extracted the data from Oracle into HDFS using Sqoop
- Using Kafka on publish-subscribe messaging as a distributed commit log, have experienced in its fast , scalable and durability .
- Test Driven Development (TDD) process and extensive experience with Agile and SCRUM programming methodology.
- Implemented POC to migrate MapReduce jobs into SparkRDD transformations using SCALA
- Scheduled map reduce jobs in production environment using Ooziescheduler .
- Involved in Cluster maintenance, Cluster Monitoring and Troubleshooting , Manage and review data backups and log files.
- Designed and implemented map reduce jobs to support distributed processing using java , Hive and ApachePig
- Analyzing Hadoopcluster and different Big Data analytic tools including Pig , Hive , Cassandra and Sqoop .
- Improved the Performance by tuning of HIVE and mapreduce .
- Research, evaluate and utilize new technologies/tools/frameworks around Hadoopecosystem
Environment : HDFS, Map Reduce, Hive, Sqoop, Pig, Flume, Vertica, Oozie Scheduler, Java, Shell Scripts, Teradata, Oracle, HBase, MongoDB, Cloudera, AWS, Javascript, JSP, Kafka, Spark, Scala and ETL, Python.
Hadoop / big data Developer
Confidential, Secaucus, NJ
Responsibilities:
- Responsible for building scalable distributed data solutions using Hadoop
- Designed the projects using MVCarchitecture providing multiple views using the same model and thereby providing efficient modularity and scalability
- Custom talend jobs to ingest , enrich and distributedata in Cloudera Hadoop ecosystem.
- Downloads the data that was generated by sensors from the Patients body activities, the data will be collected in to the HDFS system online aggregators by Kafka .
- Exploring with Spark improving the performance and optimization of the existing algorithms in Hadoop using Sparkcontext , Spark - SQL , Data Frame , pair RDD's , Spark YARN .
- Improving the performance and optimization of existing algorithms in Hadoop using Sparkcontext , Spark - SQL and SparkYARN using Scala .
- Implemented SparkCore in Scala to process data in memory.
- Performed job functions using Spark API's in Scala for realtimeanalysis and for fastqueryingpurposes .
- Spark Streaming collects this data from Kafka in near-real-time and performs necessary transformations and aggregation on the fly to build the common learner data model and persists the data in NoSQL store ( Hbase ).
- Have done OOP's and functional programming on SCALA .
- Enhanced and optimized product Spark code to aggregate , group and rundata mining tasks using the Sparkframework .
- Handled importing of data from various data sources, performed transformations using MapReduce , Spark and loaded data into HDFS .
- Developed workflow in Oozie to orchestrate a series of Pig scripts to cleanse data, such as removing personal information or merging many small files into a handful of very large, compressed files using pigpipelines in the data preparation stage.
- Used Pig in three distinct workloads like pipelines , iterativeprocessing and research .
- Used PigUDF's in Python , Java code and uses sampling of large data sets.
- Involved in moving all log files generated from various sources to HDFS for further processing through Flume and process the files by using some piggybank .
- Extensively used PIG to communicate with Hive using HCatalog and HBASE using Handlers.
- Created PIG Latin scripting and Sqoop Scripting.
- Involved in transforming data from legacy tables to HDFS , and Cassandra tables using Sqoop
- Implemented exception tracking logic using Pigscripts
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it
Environment : Hadoop, Map Reduce, Spark, shark, Kafka, HDFS, Hive, Pig, Oozie, Core Java, Eclipse, Hbase, Flume, Cloudera, Oracle 10g, UNIX Shell Scripting, Scala, MongoDB, HBase, Cassandra, Python.
Sr.java/j2ee developer:
Confidential
Responsibilities:
- Used JSF framework to implement MVC design pattern.
- Developed and coordinated complex high - quality solutions to clients using J2SE , J2EE , Servlets , JSP , HTML , Struts , Spring MVC , SOAP , JavaScript , JQuery , JSON and XML .
- Wrote JSF managed beans , converters , and validator's following framework standards and used explicit and implicit navigations for page navigations.
- Designed and developed Persistence layer components using HibernateORMtool .
- UI designed using JSF tags, ApacheTomahawk & Rich faces.
- Oracle10g used as backend to store and fetch data.
- Experienced in using IDE's like Eclipse and NetBeans , integration with Maven
- Creating Real-time Reporting systems and dashboards using xml , MySQL , and Perl
- Working on Restful web services which enforced a stateless client server and support JSON (few changes from SOAP to RESTFUL Technology) Involved in detailed analysis based on the requirement documents.
- Involved in Design, development and testing of web application and integration projects using Object Oriented technologies such as Core Java , J2EE , Struts , JSP , JDBC , Spring Framework, Hibernate , JavaBeans , WebServices ( REST / SOAP ), XML , XSLT , XSL , and Ant .
- Designing and implementing SOA compliant management and metrics infrastructure for MuleESB infrastructure utilizing the SOA management components.
- Used Node JS for server-side rendering. Implemented modules into Node JS to integrate with designs and requirements.
- JAX - WS used to interact in front-end module with backend module as they are running in two different servers.
- Responsible for Offshore deliverables and provide design/technical help to the team and review to meet the quality and time lines.
- Migrated existing Struts application to Spring MVC framework.
- Provided and implemented numerous solution ideas to improve the performance and stabilize the application.
- Extensively used LDAP Microsoft Active Directory for user authentication while login.
- Developed unit test cases using JUnit .
- Created the project from scratch using AngularJS as frontend, NodeExpressJS as backend.
- Involved in developing Perl script and some other scripts like java script
- Tomcat is the webserver used to deploy OMS web application.
- Used SOAP Lite module to communicate with different web - services based on given WSDL .
- Prepared technical reports &documentation manuals during the program development.
Environment : JDK 1.5, JSF, Hibernate 3.0, JIRA, NodeJS, Cruise control, Log4j, Tomcat, LDAP, JUNIT, NetBeans, Windows/UNIX.
Java Developer
Confidential
Responsibilities:
- Performed analysis for the client requirements based on the developed detailed design documents.
- Developed UseCases , ClassDiagrams , SequenceDiagrams and Data Models .
- Developed STRUTS forms and actions for validation of user request data and application functionality.
- Developed JSP's with STRUTS custom tags and implemented JavaScript validation of data.
- Developed programs for accessing the database using JDBC thin driver to execute queries, Prepared statements, Stored Procedures and to manipulate the data in the database
- Used JavaScript for the web page validation and Struts Validator for server - side validation
- Designing the database and coding of SQL, PL / SQL , Triggers and Views using IBMDB2 .
- Performance tuning in Datastage and sql code.
- Designed Server/Parallel jobs in Datastage Designer to replace materialized views and in this process answered the requirements alongside.
- Used DatastageDirector to monitor and schedule the jobs . • Developed Message Driven Beans for asynchronous processing of alerts. • Used ClearCase for source code control and JUNIT for unit testing. • Involved in peer code reviews and performed integration testing of the modules . Followed coding and documentationstandards . Environment : Java, Struts, JSP, JDBC, XML, Junit, Rational Rose, CVS, DB2, Windows.
