We provide IT Staff Augmentation Services!

Senior Hadoop Developer Resume

2.00/5 (Submit Your Rating)

Austin, TX

PROFESSIONAL SUMMARY:

  • Over all 9 years of extensive IT experience in all phases of Software Development Life Cycle (SDLC), which includes 4 years of strong experience working on Apache Hadoop and Spark ecosystems.
  • Worked extensively with Hadoop Distributions like Cloudera, Hortonworks and have in depth understanding of Hadoop Architecture including YARN and various components such as HDFS, Resource Manager, Node Manager, Name Node, Data Node and MR v1 & v2.
  • Having 4 years of work experience in data ingestion, storage, querying, processing and analysis of Big Data using analytics technology stack (Map Reduce, Pig, Hive, Sqoop, Oozie, Zookeeper, Flume, Spark, Spark SQL, Streaming, Yarn, Python, R, Tableau, AtScale, and Talend) with NoSQL databases such as HBase, Cassandra.
  • Executed all phases of a Big Data project life cycle during the tenure that includes Scoping Study, Requirements gathering, Design, Development, Implementation, Quality Assurance, application Support for end - to-end IT solution offerings.
  • Exclusive experience in Hadoop and its components like HDFS, Map Reduce, Apache Pig, Hive, Sqoop, HBase, and Oozie.
  • Extensive Experience in Setting Hadoop Cluster.
  • Imported/Exported data using Sqoop in between RDBMS and HDFS on regular basis.
  • Having good experience in writing Hive, Impala queries for generating reports in Tableau.
  • Designed flume agents configuration and implementing them in cluster for ingesting data from external sources into HDFS.
  • Worked on defining standards and guidelines for Hadoop and Hive.
  • Having good knowledge on Machine Learning algorithms (Supervised and Unsupervised).
  • Experience with configuration of Hadoop Ecosystem components Map Reduce, Hive, Hbase, Pig, Sqoop, Oozie, Flume, Storm, Spark, Yarn and Tez.
  • Having experience with processing real time streamed data using Spark streaming
  • Worked on Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data. Utilized Spark SQL with Data Frames to efficiently process data and use these data to feed the analytical warehouses.
  • Implemented business functionality using Spark SQL jobs in functional programming language Scala
  • Good Knowledge on Kafka and Storm.
  • Defined best practices for Tableau report development.
  • Designed and implemented proof of concept solutions and created advanced BI visualizations using Tableau platform.
  • Having good understanding of security implementation MIT Kerberos and LDAP.
  • Extensive experience in using ETL Tool Talend features such as context variables, triggers, connectors for Database and flat files.
  • Hands on Experience on many components, which are there in the palette to design Jobs & used Context Variables to Parameterize Talend Jobs.
  • Involved in performance tuning of the existing hive queries and map reduce jobs, Impala queries by analyzing query profiles.
  • Involved in in-room/telephonic Scrum meetings to gather requirements and analyzing the requirements and developments.
  • Experience in different operating Systems UNIX, Linux, and Windows.
  • Executed projects using Java/J2EE technologies such as Core Java, Servlets, Jsp, JDBC, Struts.
  • Experience in application development frameworks like spring, Hibernate and also on validation plug-ins like Validator frameworks.
  • Experience with being a technical lead of a team of engineers.
  • Strong experience with version control tools such as Subversion, AccuRev and Git.
  • Experienced in Developing J2EE Application on IDE tools like Eclipse, IntelliJ.
  • Expertise in build scripts like ANT and Maven and build automation.
  • Strong Experience in working with Databases like Oracle, SQL Server, EDB and proficiency in writing complex SQL, PL/SQL for creating tables, views, indexes, stored procedures and functions.
  • Experience in web design using HTML, Bootstrap, XML, CSS, AngularJS, NodeJS, AJAX, JavaScript, EXT JS and JQuery.
  • Good working knowledge on RESTful Web services with JSON using Jackson API.
  • Having experience in understanding of existing systems, maintenance and production support, on technologies such as Java, J2EE and various databases (Oracle, SQL Server).
  • Knowledge on Flume, NoSQL, Impala, Python, R, ETL(Talend), Tableau, Solr Search, Spark Ecosystem (Spark Core, SQL, Streaming, ML Lib, Graph X) Data Warehouse and BI technologies.
  • Good Knowledge on Cloud Suites like AWS, Azure and Machine Learning algorithms like Linear Regression, Logistic Regression, Decision Tree and Random Forest tree.
  • Experience in problem solving, analysis, implementation, installation and configuration skills.
  • Experience with all stages of the SDLC and Agile Development model right from the requirement gathering to Deployment and production support.
  • Worked on POC for accessing the same data directory by two different tables with different data types instead of loading the data explicitly from one table to other table.
  • Good interpersonal skills, commitment, result oriented, hard working with a quest and zeal to learn new technologies and undertake challenging tasks. Excellent team member with strong communication skills and capable of meeting set deadlines.

TECHNICAL SKILLS:

Languages: Java, J2EE, Scala, SQL, Python, R

Hadoop Distributions: Cloudera, Hortonworks, Apache Hadoop, MapR

Big Data Stack: Hadoop, HDFS, Map Reduce, Hive, Pig, Sqoop, Oozie, HBase, Yarn, Flume, Impala, Zookeeper, Scala, Spark Core, Spark SQL, Spark Streaming, Spark MLlib, Kafka, Storm, Tableau, ETL(Talend), Cassandra, AWS, Machine Learning (Linear and Logistic Regression), Solr search

Web Technologies: J2EE(Servlets,Jsp,JDBC), JavaScript, JQuery, AngularJs, NodeJS, Python

Web, Application Servers: JBoss, BEA Web Logic and Tomcat

Frameworks: Struts, Hadoop, Ext JS, Spring, Hibernate.

Java IDEs: Eclipse and My Eclipse, IntelliJ.

Tracking: Build Tools and Version Control Jira, Ant, Maven,SVN, GIT, AccuRev

Databases: Oracle 11g, SQL Server 2008 R2, EDB (Postgre SQL), My SQL.

No SQL: HBase, Cassandra

Operating Systems: Windows, Linux and Unix

PROFESSIONAL EXPERIENCE:

Confidential

Senior Hadoop Developer, Austin, TX

Responsibilities:

  • Used to provide services to customers for their Payment Analytics, Fraud Analytics and Risk Analytics using Hadoop as a data lake.
  • Ingested data from different sources into Hadoop Data Lake using Sqoop.
  • Created Hive external tables for each source table in Hadoop Data Lake.
  • Created Hive Partitions and Buckets for tables on basis of load date (Transaction Day).
  • Worked extensively with Hive to improve query performance.
  • Built reusable Hive UDF libraries for business requirements, which enabled users to use these UDF's in Hive Querying.
  • Worked closely with data warehouse architect and business intelligence analyst to develop solutions.
  • Developed Pig Scripts for processing the Ingested data as per business requirement.
  • Developed Pig UDF’s in Java in order to convert date formats as desired.
  • Created Oozie workflows to automate the actions required for each job.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java Map Reduce, Hive, Pig and Sqoop.
  • Worked on POC for accessing the same data directory by two different tables with different data types instead of loading the data explicitly from one table to other table.
  • Worked on different file formats (Text File, Sequence File, RCFile and Parquet) to improve the query performance and reduce the size of disk space.
  • Experience in Using Cloudera Hue Interface for Job monitoring and viewing table metadata.
  • Developed Shell Scripts for executing and parameterizing each Job.
  • Implemented Spark as a processing engine to achieve RDD using Scala.
  • Having good understanding of security implementation MIT Kerberos, Apache Sentry.
  • Actively debugging the application teams P1, P2 issues on the clusters.
  • Worked closely with MRA's to create reports/dashboards using tableau desktop
  • Reviewed basic SQL queries and altered for better performance in Tableau Desktop.
  • Created views in Tableau Desktop that was published to internal team for review and further data analysis and customization using filters and actions.
  • Designed and developed various analytical reports from multiple data sources by blending data on a single worksheet in Tableau Desktop
  • Worked on DB Visualizer to query Hive tables by using Kerberos authentication.
  • Worked on Failover Mechanism (Data Recovery using Recover Partitions/Invalidate metadata).
  • Implemented business functionality using Spark core and Spark SQL with Scala.
  • Transforming the Hive jobs into Spark SQL.
  • Created dashboards using cloudera solr search
  • Imported, exported data from various databases Oracle into HDFS using Talend.
  • Supported Business users during UAT.

Environment: Hadoop, Cloudera 5.8, HDFS, Map Reduce, Hive, Pig, Sqoop, Oozie, Solr search, Spark core, Spark SQL, Scala, SQL, Java/J2EE, Eclipse, File Formats, Talend, Tableau, ETL, At Scale.

Confidential

Senior Hadoop Developer, Austin, TX

Responsibilities:

  • Collaborate with subject matter experts, various stakeholders and fellow developers to design, develop, implement and support data analytics.
  • Created Hive tables, and loading and analyzing data using hive queries
  • Worked on Performance Tuning of Hadoop jobs by applying techniques such as Partitioning, Bucketing and using different file formats such as Parquet and ORC.
  • Defined job work flows as per their dependencies in Oozie.
  • Involved in creating the Oozie workflows to run multiple Hive and Pig jobs, which run independently with time and data availability.
  • Used JAVA, J2EE application development skills with Object Oriented Analysis and extensively involved throughout Software Development Life Cycle (SDLC)
  • Proactively monitored systems and services, architecture design and implementation of Hadoop deployment, configuration management, backup, and disaster recovery systems and procedures.
  • Used Flume to collect, aggregate, and store the web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
  • Lead the team during data ingestion from external system.
  • Played challenging roles such as Lead member of the team, Onsite/Offshore coordinator and Subject Matter Expert (SME).
  • Collaborated with application owners and other engineers to define requirements to design, build and tune complex solutions.
  • Load and transform large sets of structured, semi structured and unstructured data
  • Supported Map Reduce Programs those are running on the cluster
  • Importing and exporting data into HDFS and Hive using Sqoop, Spark Core and Spark SQL.
  • Worked on converting Hive queries into Spark transformations using Spark RDDs and SQL.
  • Analyze Performance of queries running on Impala.
  • Worked on defining standards and guidelines for Hadoop and Hive.
  • Developed Scala and SQL code to extract data from various databases.
  • Developed Spark code using Scala and Spark-SQL for faster testing and data processing.
  • Managed and reviewed log files
  • Used to create indexes and dashboard using Solr Search
  • Having good amount of knowledge on Cassandra and HBase.
  • Implemented partitioning, dynamic partitions and buckets in HIVE

Environment: Hadoop, Cloudera 5.8, HDFS, Map Reduce, Hive, Pig, Sqoop, Oozie, Flume, Spark Core, Spark SQL, Scala, Oracle, Java/J2EE, Eclipse, File Formats (RC, ORC, Parquet), Tableau Desktop.

Confidential

Hadoop Developer

Responsibilities:

  • Worked on data ingestion from Oracle to HDFS and vice-versa using SQOOP.
  • Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
  • Implemented Hive tables and HQL Queries for the reports
  • Worked with different file formats and compression techniques to determine standards importing and exporting data into HDFS and Hive using Sqoop
  • Developed Hive queries to analyze/transform the data in HDFS.
  • Designed and Implemented Partitioning (Multi-level), Buckets in HIVE.
  • Designed and implemented pig UDFs for evaluation, filtering, loading and storing of data.
  • Analyzing/Transforming data with Hive and Pig.
  • Resolved issues with flume and advised flume Agent configuration changes.
  • Worked on debugging, performance tuning of Hive & Pig Jobs.
  • Developed pig scripts to preprocess data before moving the data into final tables
  • Developed Pig scripts to process the data from different data sets and generating the aggregating results.
  • Creating Views on Hive tables.
  • Effective coordination with offshore team and managed project deliverable on time.
  • Worked on QA support activities, test data creation and Unit testing activities.
  • Responsible for creating Hive tables, loading the structured data resulted from Map Reduce jobs into the tables and writing hive queries to further analyze the logs to identify issues and behavioral patterns.
  • Used Hive to analyses data ingested into Hbase by using Hive-Hbase integration and compute various metrics for reporting on the dashboard.
  • Developed job flows to automate the workflow for pig and hive jobs.
  • Extensively involved in performance tuning of Oracle queries.
  • Elastic search performance and optimization
  • Extensively worked in agile environment, with daily scrum meeting, stand up meetings, presentations and review
  • Implemented monitoring and established best practices around usage of elastic search
  • Loaded the aggregated data onto Oracle from Hadoop environment using Sqoop for reporting on the dashboard.
  • Written UNIX scripts to automate batch functions.
  • Used Tableau for visualization on processed data.

Environment: Hadoop Ecosystem, Apache Pig, Hive, Sqoop, Oozie, Tableau, HCat, Java, UNIX Scripts, Oracle, Cent OS, Map Reduce, Hbase, Cassandra, Elastic Search.

Confidential

Software Engineer

Responsibilities:

  • Involved in Use Case meeting to understand and analyze the requirements, Coded as per Prototype.
  • Developed various UI (User Interface) components using Struts (MVC), JSP, and HTML.
  • Developed Controllers, created JSPs and configured in Struts-config.xml, Web.xml files.
  • Developed MVC architecture, Business Delegate, Service Locator, Session facade, and Data Access Object and Singleton patterns
  • Involved in writing all client side validations using Java Script, JSON.
  • Involved in the complete development, testing and maintenance process of the application.
  • Used Hibernate 2.0 as the ORM tool to communicate with the database.
  • Designed and created a web-based test client using Struts up on client’s request, which is used to test the different parts of the application.
  • Involved in writing the test cases for the application using JUnit.
  • Used extensive JSP, HTML, and CSS to develop presentation layer to make it more user friendly.
  • Involved in different Testing phases like Unit Test, Integration Test and Regression Test.
  • Involved in Development process and have knowledge in usage of Tracker Tools like JIRA.
  • Developed back-end stored procedures and triggers using Oracle PL/SQL, involved in database objects creation, performance tuning of stored procedures, and query plan
  • Responsible for developing and maintaining all the session beans.
  • Supported the application through debugging, fixing, production support and maintenance releases.
  • Worked closely with the client and the offsite team; coordinated activities between them for effective implementation of the project.
  • Developed web services using RESTful
  • Involved in Performance tuning of Application
  • Involved in Restful Web services with JQuery using Jackson API,
  • Involved in Web services (SOAP, RESTful) Testing using Confidential EAM Web Service tool kit.

Environment: Core Java, JSP, Servlets, Struts, EJB2.0, Ext JS, XML, Oracle 11g, PostgreSQL, Java Script, Web Service, SQL Server 2008R2, Eclipse, TOAD, JIRA, SVN, Tortoise, Log4j.

Confidential

Software Engineer

Responsibilities:

  • Involved in Use Case meeting to understand and analyze the requirements, Coded as per Prototype.
  • Involved in product development and customizations using Confidential CRB studio Designing and development of BIO’s, record sets, data sources, forms, shells, navigations
  • Involved in Coding and bug fixing.
  • Used Validate plug-in of struts framework to handle server side validations.
  • Involved in usage of SVN for version control process.
  • Logging of errors in application is achieved by using Log4J API.
  • Tables are created by running Scripts on SQL Developer IDE.
  • Involved in Build, Deployment Activities and Oracle/SQL Database Schema Restore and backup.
  • Involved in resolving Memory Leakage and JVM Problems.
  • Involved in Unit Test, Integration Test and Regression Test plans.
  • Functional and UI design has been prepared
  • Implementation at BIO level
  • Creation of Record sets and BIOs for the database schema
  • Created Relationships for data Integrity
  • Created Lookups and attribute domains
  • Implementation at UI level
  • Menus for Navigation
  • Forms for various Perspectives
  • Implemented shells like List Shell, Detail Shell, Tab Group Shell, Toggle Shell to
  • Provide better look and feel
  • Toolbars to allow UI Actions for Buttons
  • Used Form Slots by considering the BIO schema
  • Authentication and authorization has been achieved by creating users and profiles in platadmin.
  • Implemented object-permissions at widget, menu, and form levels
  • Developed Form level extensions to achieve UI level validations and BIO level extensions to fulfill Functional requirements and validations.
  • All required data is entered by using Bulk Import.
  • Involved in Development process and have knowledge in usage of Tracker Tools like JIRA.
  • Having good Knowledge in Epiphany Platform (Open Architecture).
  • Having Extensive Hands on Experience on Complex PL/SQL Programming.

Environment: s: CRB Studio, Web logic server 8.1, LDAP, Core Java, Servlets, Struts, Spring, Hibernate, Java Beans, Db2, SQL Server, Eclipse, Oracle, XML, Apache Tomcat, SVN, Windows, UNIX.

Confidential

Associate Software Engineer

Responsibilities:

  • Involved in translating Business Requirement into Technical Requirement.
  • Involved in all the phases of SDLC including Requirements Collection, Design and analysis of the customer specification from the Business Analyst.
  • Used Java & J2EE frameworks to build the application while storing the data and retrieving from different data sources like Oracle and SQL Server.
  • Wrote the scripts to support different environments like Windows and Linux.
  • Involved in writing database object creation and stored procedures in Oracle and SQL Server and Involved in Code review as well.
  • Used JDBC API to communicate with the database.
  • Client-side validation was done using JavaScript.
  • Extensively involved in the development of Java Server Pages (JSP).
  • Involved in performance and SQL Query optimization.
  • Reviewing, sending the Status Report consisting of the Bug reports, defect logs and the other deliverables along with the summary of the entire work done on a weekly basis to the Client & Project Manager.
  • Extensively used Jira, SVN and eclipse.

Environment: Core Java, J2EE, JDBC, JSP, SQL, Eclipse, Oracle, XML, Apache Tomcat, SVN, JUnit, Windows, UNIX.

We'd love your feedback!