Sr. Hadoop/spark Developer Resume
West Des Moines, IA
PROFESSIONAL SUMMARY:
- Having 9+ years of experience in Analysis, Design, Development, Integration, Testing and maintenance of various applications using JAVA /J2EE technologies along with Around 4 years of Big Data /Hadoop experience.
- Knowledge in building highly scalable Big - data solutions using Hadoop and multiple distributions i.e. Cloudera, Horton works and NoSQL platforms (HBase & Cassandra).
- Good Knowledge in big data architecture with Hadoop File system and its eco system tools Map Reduce, HBase, Hive, Pig, Zookeeper, Oozie, Flume, Avro, Impala and Apache spark.
- Knowledge on performing Data Quality checks on petabytes of data Solid understanding of Hadoop MRV1 and Hadoop MRV2 (or) YARN Architecture.
- Good knowledge on Amazon AWS concepts like EMR & EC2 web services which provides fast and efficient processing of Big Data.
- Developed, deployed and supported several Map Reduce applications in Java to handle semi and unstructured data.
- Experience in writing Map Reduce programs and using Apache Hadoop API for analyzing the data.
- Strong experience in developing, debugging and tuning Map Reduce jobs in Hadoop environment.
- Expertise in developing PIG and HIVE scripts for data analysis
- Hands on experience in data mining process, implementing complex business logic and optimizing the query using Hive QL and controlling the data distribution by partitioning and bucketing techniques to enhance performance.
- Experience working with Hive data, extending the Hive library using custom UDF's to query data in non- standard formats.
- Experience in performance tuning of Map Reduce, Pig jobs and Hive queries
- Involved in the Ingestion of data from various Databases like TERADATA (Sales Data Warehouse), AS400, DB2 and SQL-SERVER using Sqoop.
- Experience working with Flume to handle large volume of streaming data.
- Extensive experience in migrating ETL operations into HDFS systems using Pig Scripts.
- Good knowledge in evaluating big data analytics libraries (MLlib) and use of Spark-SQL for data exploratory.
- Good Knowledge in implementing advanced procedures like text analytics and processing using the in-memory computing capabilities like Apache Spark written in Scala.
- Experience in implementing a distributed messaging queue to integrate with Cassandra using Apache Kafka and zookeeper.
- Expert in creating and designing data ingest pipelines using technologies such as Spring Integration, Apache Storm-Kafka.
- Experience with Oozie Workflow Engine in running workflow jobs with actions that run Hadoop Map Reduce and Pig jobs.
- Developed core modules in large cross-platform applications using JAVA, J2EE, Hibernate, Python, Spring, JSP, Servlets, EJB, JDBC, JavaScript, XML, and HTML.
- Experienced with build tools Maven, ANT and continuous integrations like Jenkins.
- Working Knowledge in configuring and monitoring tools like Ganglia and Nagios.
- Hands-on experience in using relational databases like Oracle, MySQL, PostgreSQL and MS-SQL Server.
- Extensive experience in developing and deploying applications using Web Logic, Apache Tomcat and JBOSS.
- Developed Unit test cases using Junit, Easy Mock and MRUnit testing frameworks.
- Experienced with version controller systems like SVN, Clear case.
- Experience using IDEs tools Eclipse, My Eclipse, RAD and NetBeans.
- Hands on development experience with RDBMS, including writing SQL queries, PLSQL, views, stored procedure, triggers, etc.
- Expertise in Waterfall and Agile software development model & project planning using Microsoft Project Planner and JIRA.
TECHNICAL SKILLS:
Big Data Technologies: HDFS, Map Reduce, Hive, Hcat, Pig, Sqoop, Flume, Oozie, Avro, Hadoop Streaming, Zookeeper, Kafka, Impala, Apache Spark, hue, Ambari. Apache ignite.
Hadoop Distributions: Cloudera (CDH4/CDH5),Horton Works
Languages: Java, C, SQL, PYTHON,PL/SQL,PIG-Latin, HQL
IDE Tools: Eclipse, NetBeans, RAD
Framework: Hibernate, Spring, Struts, Junit
Web Technologies: HTML5, CSS3, JavaScript, JQuery, AJAX, Servlets, JSP,JSON, XML, XHTML, JSF, Angular JS
Web Services: SOAP,REST, WSDL, JAXB, and JAXP
Operating Systems: Windows (XP,7,8), UNIX, LINUX, Ubuntu, CentOS
Application Servers: JBoss, Tomcat, Web Logic, Web Sphere
Reporting Tools: /ETL Tools Tableau, Power view for Microsoft Excel, Informatica
Oracle, MySQL, DB2, Derby, PostgreSQL, No: SQL Database (HBase, Cassandra)
PROFESSIONAL EXPERIENCE:
Confidential, West Des Moines, IA
Sr. Hadoop/Spark Developer
Responsibilities:
- Good Knowledge in designing and deployment of Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, Oozie, Sqoop, Kafka, Spark, Impala with Cloudera distribution.
- Worked on Cloudera distribution and deployed on AWS EC2 Instances.
- Knowledge on experience on Cloudera Hue to import data on the GUI.
- Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing.
- Performed Data Ingestion from multiple internal clients using Apache Kafka.
- Implemented real time system with Kafka, Storm and Zookeeper .
- Worked on integrating Apache Kafka with Spark Streaming process to consume data from external REST APIs and run custom functions.
- Developed and Configured Kafka brokers to pipeline server logs data into Spark streaming.
- Involved in performance tuning of Spark jobs using Cache and using complete advantage of cluster environment.
- Developed Spark scripts by using Scala Shell commands as per the requirement.
- Configured, deployed and maintained multi-node Dev and Tested Kafka Clusters.
- Configured spark streaming data to receive real time data from Kafka and store it in HDFS.
- Data from HDFS, Hive and HBase for search and quick BI visualization in Kibana .
- Built centralized logging to enable better debugging using Elastic Search , Logstash and Kibana .
- Efficiently handled periodic exporting of SQL data into Elastic search .
- Developed in scheduling Oozie workflow engine to run multiple Hive and Pig jobs.
- Involved in running Hadoop streaming jobs to process terabytes of text data. Worked with different file formats such as Text, Sequence files, Avro, ORC and Parquet.
- Configured, supported and maintained all network, firewall, storage, load balancers, operating systems, and software in AWS EC2.
- Implemented the use of Amazon EMR for Big Data processing among a Hadoop Cluster of virtual servers on Amazon related EC2 and S3.
- Developed PIG Latin scripts to extract the data from the web server output files and to load into HDFS.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Worked on creating Hive tables and written Hive queries for data analysis to meet business requirements and experienced in Sqoop to import and export the data from Oracle & MySQL.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs.
- Good knowledge in using Data Manipulations, Tombstones, Compactions in Cassandra. Well experienced in avoiding faulty Writes and Reads in Cassandra.
- Involved in creating data-models for customer data using Cassandra Query Language related to Cassandra clusters.
- Implemented YARN Capacity Scheduler on various environments and tuned configurations according to the application wise job loads.
- Configured Continuous Integration system to execute suites of automated test on desired frequencies using Jenkins, Maven & GIT.
- Involved in loading data from LINUX file system to HDFS .
- Followed Agile Methodologies while working on the project.
Environment: Hadoop, HDFS, Hive, Spark, Cloudera, AWS EC2, S3, EMR, Sqoop, Kafka, Yarn, Shell Scripting, Impala, Scala, Pig, Cassandra, Oozie, Java, JUnit, Agile methods, Jenkins, Maven, Linux, MySQL, Elastic Search, Kibana
Confidential, NY
Hadoop Developer
Responsibilities:
- Working in an Agile team to deliver and support required business objectives by using Java, Python and shell scripting and other related technologies to acquire, ingest, transform and publish data both to and from Hadoop Ecosystem.
- Loaded data into the cluster from dynamically generated files using Flume and from relational database management systems using Sqoop.
- Worked on large-scale Hadoop YARN cluster for distributed data processing and analysis using Connectors, Spark core, Spark SQL, Sqoop, Pig, Hive and NoSQL databases.
- Used Flume to collect, aggregate and store the web log data onto HDFS.
- Performed Data Cleansing using Python and loaded into the target tables.
- Logical implementation and interaction with HBASE.
- Used Scala to store streaming data to HDFS and to implement Spark for faster processing of data.
- Integrating user data from Cassandra to HDFS. Integrating Cassandra with Storm for real time user attributes look up.
- Performed Sqoop Incremental imports by using Oozie based on every day.
- Installed and configured Hadoop MapReduce, HDFS, developed MapReduce jobs in Java for data cleaning and pre-processing.
- Involved in using HCATALOG to access Hive table metadata from MapReduce or Pig code.
- Created Pig scripts to transform the HDFS data and loaded the data into Hive external table.
- Implemented Spark Scripts using Scala, Spark SQL to access hive tables into spark for faster processing of data.
- Performed Optimizations of Hive Queries using Map side joins, dynamic partitions and Bucketing.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
- Performed advanced procedure like text analytics and processing, using the in-memory computing capabilities of Spark using Scala.
- Implemented Spark RDD transformations, actions to implement business analysis.
- Connected to HDFS using Pentaho Kettle to read data from hive tables and perform analysis.
- Worked on Spark Streaming and Spark SQL to run sophisticated applications on Hadoop.
- Used Oozie and Oozie coordinators to deploy end to end processing pipelines and scheduling the work flows.
- Worked on concept of quorum with Kafka and Zookeeper.
- Deployed reports on Pentaho BI Server to give central web access to the users.
- Created several dashboards in Pentaho using Pentaho Dashboard Designer.
- Created and maintained Technical documentation for launching Hadoop Clusters and executing pig Script.
Environment: Hadoop, CDH, Scala, MapReduce, HDFS, Hive, Pig, Sqoop, HBASE, Flume, Spark SQL, Spark-Streaming, MapR, Pentaho, Python, UNIX Shell Scripting and Cassandra.
Confidential, Salt lake city, Utah
Hadoop Developer
Responsibilities:
- Worked on analyzing Hadoop cluster using different big data analytic tools including Pig, Hive, and MapReduce.
- Responsible to managing data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data.
- Used Pig as ETL tool to do transformations, event joins and some pre-aggregations before storing the data into HDFS.
- Worked with Zookeeper, Oozie, and Data Pipeline Operational Services for coordinating the cluster and scheduling workflows.
- Involved in cluster maintenance which includes adding, removing cluster nodes, cluster monitoring and troubleshooting, reviewing and managing data backups and Hadoop log files.
- Used HIVE join queries to join multiple tables of a source system and load them into Elasticsearch Tables.
- Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
- Collected and aggregated large amounts of web log data from different sources such as web servers, mobile and network devices using Apache Flume and stored the data into HDFS for analysis.
- Designed and Modified Database tables and used HBase Queries to insert and fetch data from tables.
- Involved in moving all log files generated from various sources to HDFS for further processing through Flume.
- Created HBase tables to store variable data formats coming from different portfolios Performed real time analytics on HBase using Java API and Rest API.
- Participated in development/implementation of Cloudera Hadoop environment.
- Worked extensively with Sqoop for importing and exporting the data from HDFS to Relational Database system and vice-versa. Loading data into HDFS.
- Involved in Setup and benchmark of Hadoop /HBase clusters for internal use.
- Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Involved in importing and exporting the data from RDBMS to HDFS and vice versa using Sqoop.
- Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Involved in running MapReduce jobs for processing millions of records.
- Involved in full life-cycle of the project from Design, Analysis, logical and physical architecture Modeling, Development, Implementation and Testing.
Environment: Hadoop, Big Data, HDFS, MapReduce, Sqoop, Oozie, Pig, Hive, HBase, Flume, LINUX, Java, Eclipse, Cassandra, Hadoop Distribution of Cloudera., PL/SQL, Windows, UNIX Shell Scripting, and Eclipse.
Confidential
Java/J2EE Developer
Responsibilities:
- Involved in various phases of Software Development Life Cycle (SDLC) of the application like Requirement gathering, Design, Analysis and Code development.
- Developed business functionalities using JAVA/J2EE, Spring, Hibernate, RESTful Web Services, Oracle database and Web logic application server.
- Developed and designed the front end using HTML, CSS and JavaScript and Ajax.
- Developed the entire application implementing MVC Architecture integrating Hibernate and Spring frameworks.
- Involved in development of presentation layer using JSP and Servlets with Development tool Eclipse IDE.
- Worked on development of Hibernate, including mapping files, configuration file and classes to interact with the database.
- Used XML/XSL and Parsing using both SAX and DOM parsers.
- Designed and developed EJBs to handle business logic and store persistent data.
- Planned and implemented various SQL, Stored Procedure and triggers.
- Written JUnit Test cases for performing unit testing.
- Responsible for testing, debugging, bug fixing and documentation of the system.
- Application deployment is done in Tomcat server.
- Implemented Single Sign-On (SSO) functionality for the application.
- Developed database stored procedures, functions and packages using SQL, PL/SQL and SQL Developer.
- Developed unit test cases using JUnit and performed Unit Testing of application modules.
- Supported QA and UAT testing. Fixed the bugs/issues reported during testing phase.
- Provided the technical guidance to new and/or junior team members.
Environment: Java, Servlets, JSP, JQuery, JUnit, Hibernate, JPA, Spring, Ajax, Oracle, JMS, Eclipse Apache Ant, Tomcat, Web Services, Apache Axis, WebSphere, JavaScript, HTML, CSS, XML, Clear Case.
Confidential
Java Developer
Responsibilities:
- Involved in various SDLC phases like Design, Development and Testing.
- Involved in developing JSP pages using Struts custom tags, jQuery and Tiles Framework.
- Developed web pages using HTML, JavaScript, JQuery and CSS.
- Used various Core Java concepts such as Exception Handling, Collection APIs to implement various features and enhancements.
- Developed server side components servlets for the application.
- Involved in coding, maintaining, and administering Servlets and JSP components to be deployed on a Web Sphere application server.
- Used automated test scripts and tools to test the application in various phases. Coordinated with Quality Control teams to fix issues that were identified
- Implemented Hibernate ORM to Map relational data directly to java objects
- Worked with Complex SQL queries, Functions and Stored Procedures.
- Involved in developing spring web MVC framework for portals application.
- Implemented the logging mechanism using log4j framework.
- Developed REST API, Web Services.
- Wrote test cases in JUnit for unit testing of classes.
- Used Maven to build the J2EE application.
- Used SVN to track and maintain the different version of the application.
- Involved in maintenance of different applications with onshore team.
- Good working experience in Tepestry processing claims.
- Working experience with professional billing claims.
Environment: Java, Spring Framework, Struts, Hibernate, RAD, SVN, Maven, Web Sphere Application Server, Web Services, Oracle Database 11g, IBM MQ, JMS, HTML, Java script, XML, CSS, REST API.
