We provide IT Staff Augmentation Services!

Hadoop Lead Developer Resume

5.00/5 (Submit Your Rating)

Dallas, TX

SUMMARY

  • Over 8+ years of professional IT experience, 4+ years in Big Data Ecosystem experience in ingestion, querying, processing and analysis of big data.
  • Excellent understanding / knowledge of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
  • Experience with configuration ofHadoopEcosystem components: Hive, HBase, Pig, Sqoop,Mahout, Zookeeper, Flume, Storm, Spark
  • Hands on experience in installing and configuring Hadoop ecosystem components like Oozie, Hive, Sqoop, Zookeeper, Pig, and Flume.
  • Good Exposure on Map Reduce programming(JAVA), Hive, Pig scripting, Spark SQL(Scala/python).
  • Experience in writing custom UDFs in java for Hive and Pig to extend the functionality.
  • Experience in writing MAPREDUCE programs in java for data cleansing and preprocessing.
  • Hands On experience on Spark, Spark Streaming,Spark Mlib, SCALA.
  • Creating the Data Frames handle inSpark with Scala.
  • Hands On experience on developing UDF, DATA Frames and SQL Queries inSpark SQL.
  • Experience in building Data pipelines using Kafka and Spark.
  • Experience in managing and reviewing Hadoop log files.
  • Hands on experience in Import/Export of data using Hadoop Data Management tool Sqoop.
  • Experience with distributed systems, large - scale non-relational data stores, RDBMS, NoSQL, map-reduce systems, data modeling, database performance, and multi-terabyte data warehouses.
  • Experience in designing, developing and implementing connectivity products that allow efficient exchange of data between our core database engine and the Hadoop ecosystem.Worked extensively on building Rapid development Framework using Core Java.
  • Extensive experience and actively involved in Requirement gathering, Analysis, Design, Reviews, Coding, Code Reviews, Unit and Integration Testing.
  • Extensive experience in designing front end interfaces using HTML, JSP, CSS, Java Script and Ajax.
  • Very familiar withdataarchitecture,Hadoopinformation architecture,datamodelingand datamining, machine learning and advanceddataprocessing.
  • Good Experience using Object Relational Mapping tool like Hibernate.
  • Experience in Spring Framework such as Spring IOC, Spring Resources, Spring JDBC.
  • Experience with various IDEs like IntelliJ, Eclipse, JBuilder and Velocity Studio.
  • Implemented the service projects on Agile Methodology and involved in running the scrum meetings.
  • Implemented the core product projects on Lean and Kanban Methodology and involved in delivering high quality health care product.
  • Experience in developing web-services using REST, SOAP, WSDL and Apache AXIS2.
  • Experience in writing the SQL queries.
  • Experience in designing and developing UI Screens using Java Server Pages, Html, CSS and JavaScript.
  • Set upSolrfor distributing indexing and search.
  • Wrote Java code to format XML documents; upload them toSolr server for indexing.
  • Loaded and accessed the process event failures messages through kafka-Solr-writer and querying onSolrcollection database.
  • Created theSolr collection in order to load the reports from the processed data. So developedSolr writer to write the encoded data toSolr collection from HDFS.
  • Familiar with data mining tools like ApacheMahoutand WEKA tools.
  • Used CVS, Maven, and SVN for Source code version control.
  • Experience in designing transaction processing systems deployed on various application servers including Tomcat, Web Sphere, Web logic.
  • Good Experience on Quality Control, JIRA, Fish Eye for tracking the tickets like accepting the tickets/defects, Submitting the tickets, Reviewing Code and closing the tickets etc.,
  • Designed dynamic user interfaces using AJAX and JQuery to retrieve data without reloading the page and send asynchronous request.
  • Excellent Experience in Code Refactoring.
  • Excellent Client interaction skills and proven experience in working independently as well as in a team.
  • Excellent communication, analytical, interpersonal and presentation skills.

TECHNICAL SKILLS

Hadoop Ecosystem: Kafka,HDFS,MapReduce,Hive,Impala,Pig,Sqoop,Flume,Oozie,Zookeeper,Ambari,Hue,Spark,Strom,Ganglia

Hadoop Platforms: Cloudera, hortonworks, MapR

Web Technologies: JDBC, Servlets, JSP, JSTL, JNDI, XML, HTML, CSS and AJAX

NoSQL Databases: HBase, Cassandra, MongoDB

Databases: Oracle 8i/9i/10g, MySQL

Languages: Java, SQL, PL/SQL, Ruby, Shell Scripting

Operating Systems: UNIX(OSX, Solaris), Windows, Linux(Cent OS, Fedora, Red Hat)

Frame Works: Struts, Hibernate, Spring, ConceptWave, ATG 7.0

Application Server: Apache Tomcat

Streaming technologies: Flume, Storm, Spark Streaming

Analytics: Spark,Mahout

Search: Elasticsearch, Solr

Application Server: Apache Tomcat

Project Management / Tools / Applications: All MS Office suites(incl. 2003), MS Exchange & Outlook, Lotus Domino Notes, Citrix Client, SharePoint, MS Internet Explorer, Firefox, Chrome, Apache, IIS

PROFESSIONAL EXPERIENCE

Confidential - Dallas, TX

Hadoop Lead Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop.
  • Involved in loading data from LINUX file system to Hadoop Distributed File System.
  • Created Hbase tables to store various data formats of PII data coming from different portfolios.
  • Experience in managing and reviewing Hadoop log files.
  • Creating instances in openstack for setting up the environment.
  • Setting up the ELK( ElatsticSearch, Logstash, Kibana) Cluster.
  • Trouble shooting any Nova, Glance issue in openstack, Kafka, Rabbitmq bus.
  • Performance testing of the environment- Creating python script to load on IO, CPU.
  • Experience with OpenStack Cloud Platform.
  • Experienced in Provisioning Hosts with flavors GP(General-purpose), SO(Storage Optimize), MO(Memory Optimize), CO(Compute Optimize).
  • Performance testing of the environment- Creating python script to load on IO, CPU.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Involved in Design, development, implementation and documentation in various big data technologies.
  • Django Framework used in developing web applications to implement the MVC architecture
  • Used Django APIs for database access
  • Design and Development of adapters to inject and eject data from various data source to/from Kafka.
  • Design and development of HBase tables according to various needs of the tenants while taking into consideration various issues related to performance.
  • Documenting the process of designing and developing the HBase tables.
  • Interact with various business teams to document the requirements for HBase tables.
  • Developed Spark applications to move data into HBase tables from various sources like Relational Database or Hive.
  • Managed and helped a team of two developers with Spark programming.
  • Design, Development and Documentation of various sqoop scripts to pull the data into Hadoop eco system.
  • Involved in creating Hive tables, and loading and analyzing data using hive queries.
  • Optimizing of existing algorithms inHadoopusingSparkContext,Spark-Sql, Data Frames and Pair RDD's.
  • Worked on Cluster of size 135 nodes.
  • ImplementedSpark using Scala and utilizing Data frames andSpark SQL API for faster processing of data
  • Worked on migrating MapReduce programs into Spark transformations usingScala.
  • Worked on reading multiple data formats on HDFS usingScala
  • Used Spark to create API's inScala for Big data analysis.
  • Created RDD's, Data Frames and Datasets.
  • Good experience with TalendOpen Studiofor designing ETL Jobs for Processing of data.
  • Used ORC, Parquet file formats for storing the data.
  • Used java code for SqlQueries and also code to retrieve the Sql Queries through Text File.
  • Used Eclipse for the Development, Testing and Debugging of the application.
  • Use python for writing script to move the data cluster to cluster.
  • Log4j framework has been used for logging debug, info & error data.
  • Created Hive External and Managed tables.
  • Use of MAVEN for dependency management and structure of the project.
  • Designed and Maintained Tez workflows to manage the flow of jobs in the cluster
  • Loaded theSpark RDD and do in memory data Computation to generate the Output response.

Environment: Spark RDD, Spark Sql, Spark Data Frames, Maven Eclipse, ElasticSearch, Logstash, Ansible, Rhel7, python, Kafka, streamsets, Influxdb, sensu, rabbitmq, Uchiwa, kibana, Hive,Pig,Hbase, Sqoop., Kibana

Confidential - Irving, TX

Hadoop Lead Developer

Responsibilities:

  • Implemented CDH3 Hadoop cluster on CentOS.
  • Implemented POC's to configure data tax Cassandra with Hadoop.
  • Launching Amazon EC2 Cloud Instances using Amazon Images (Linux/Ubuntu) and Configuring
  • Installed the application on AWS EC2 instances and configured the storage on S3 buckets.
  • Responsible for migrating the code base from Cloudera Platform to Amazon EMR and evaluated Amazon eco systems components like Redshift, Dynamo DB.
  • POC on Data Search using Elastic Search
  • Along with the Infrastructure team, involved in design and developed Kafka and Storm based data pipeline. This pipeline is also involved in Amazon Web Services EMR, S3 and RDS.
  • Worked onMongoDBby using CRUD (Create, Read, Update and Delete), Indexing, Replication and Sharding features.
  • Launched instances with respect to specific applications.
  • Developed and managed cloud VMs with AWS EC2 command line clients and management console.
  • Lead initiatives in developing cloud-based, SaaS solutions for design market.
  • Launching and Setup of HADOOP Cluster which includes configuring different components of HADOOP.
  • Import the data from different sources like HDFS/HBase into Spark RDD
  • Developed RDD's/Data Frames in Spark using Scala and Python and applied several transformation logics to load data from Hadoop Data Lake to Cassandra DB.
  • Involved in developing Map-reduce framework, writing queries scheduling map-reduce
  • Involved in Performance Optimization of Queries & Stored Procedures by analyzing Query Plans, blocking queries, Identifying missing indexes
  • Real time streaming of data using Spark with Kafka.
  • Created tables in HIVE and after that load data from HDFS to HIVE
  • Hands on experience in loading data from UNIX file system to HDFS.
  • Experienced with Performing Cassandra Query operations using Thrift API to perform real time analytics.
  • Worked on custom Pig Loaders and Storage classes to work with a variety of data formats such as JSON, Compressed CSV, etc.
  • Cluster coordination services through Zookeeper.
  • Implemented Persistence and search of data using Solr.
  • Installed and configured Flume, Hive, Pig, Sqoop and Oozie on the Hadoop cluster.
  • Involved in creating Hive tables, loading data and running hive queries in those data.
  • Extensive Working knowledge of partitioned table, UDFs, performance tuning, compression-related properties, thrift server in Hive.
  • Experience in developing Maven and ANT scripts to automate the compilation, deployment and testing of web application
  • Developed Apache Spark jobs using Scala in test environment for faster data processing and used Spark SQL for querying
  • Involved in writing optimized Pig Script along with involved in developing and testing Pig Latin Scripts.
  • Working knowledge in writing Pig's Load and Store functions.

Environment: Apache Hadoop, MapReduce, Scala, HDFS, Python, CentOS, Zookeeper, Sqoop, Kafka, MemSQL, Cassandra, Redshift, Dynamo DB, Solr, Hive, Pig, Oozie, Spark SQL, Scala, GitHub, Json, Netezza, Cassandra, Cloudera CDH3, Oracle, Maven, Ant, Eclipse, Amazon EC2, MongoDB, EMR, S3.

Confidential - Raleigh, NC

Hadoop Developer

Responsibilities:

  • Installed and configured Pig and also written Pig Latin scripts.
  • Involved in managing and reviewing Hadoop Job tracker log files and control-m log files.
  • Scheduling and managing cron jobs, wrote shell scripts to generate alerts.
  • Monitoring and managing daily jobs, processing around 200k files per day and monitoring those through RabbitMQ and Apache Dashboard application.
  • Used Control-m scheduling tool to schedule daily jobs.
  • Experience in administering and maintaining a Multi-rack Cassandra cluster
  • Monitored workload, job performance and capacity planning using InsightIQ storage performance monitoring and storage analytics, experienced in defining job flows.
  • Got good experience with NOSQL databases like Cassandra, Hbase.
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
  • Used Sqoop to efficiently transfer data between databases and HDFS and used Flume to stream the log data from servers/sensors
  • Developed MapReduce programs to cleanse the data in HDFS obtained from heterogeneous data sources to make it suitable for ingestion into Hive schema for analysis.
  • Used Hive data warehouse tool to analyze the unified historic data in HDFS to identify issues and behavioral patterns.
  • The Hive tables created as per requirement were internal or external tables defined with appropriate static and dynamic partitions, intended for efficiency.
  • Worked on setting up High Availability for GPHD 2.2 with Zookeeper and quorum journal nodes.
  • Used Control-m scheduling tool to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce, Hive and Sqoop as well as system specific jobs
  • Worked with BI teams in generating the reports and designing ETL workflows on Tableau.
  • Involved in Scrum calls, Grooming and Demo meeting, Very good experience with agile methodology.

Environment: Apache Hadoop 2.3, gphd-1.2, gphd-2.2, Map Reduce 2.3, HDFS, Hive, Java 1.6 & 1.7, Cassandra, Pig, SpringXD, Linux, Eclipse, RabbitMQ, Zookeeper, PostgresDB, Apache Solar, Control-M, Redis., Tableau, Qlikview, DataStax.

Confidential - Charlotte, NC

Hadoop Developer

Responsibilities:

  • Installed and configured Hadoop Map reduce, HDFS, Developed multiple Map Reduce jobs in java for data cleaning and preprocessing.
  • Installed and configured Pig and also written Pig Latin scripts.
  • Developed PIG scripts using Pig Latin.
  • Involved in managing and reviewing Hadoop log files.
  • Exported data using Sqoop from HDFS to Teradata on regular basis.
  • Developing Scripts and Batch Job to schedule various Hadoop Program.
  • Written Hive queries for data analysis to meet the business requirements.
  • Creating Hive tables and working on them using Hive QL.
  • Experienced in defining job flows.
  • Got good experience with NOSQL databases like Cassandra.
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
  • Designed and implemented Map reduce-based large-scale parallel relation-learning system
  • Setup and benchmarked Hadoop clusters for internal use.
  • Worked with BI teams in generating the reports and designing ETL workflows on Tableau.
  • Monitoring the log flow from LM Proxy to ES-Head.
  • Used secportal as front end of Gracie where we perform the search operations.
  • Wrote the Map Reduce code for the flow from Hadoop Flume to ES Head.

Environment: Cloudera Hadoop(CDH 4.4), Map Reduce, HDFS, Hive, Java, Pig, Cassandra, Linux, XML, MySQL, MySQL Workbench, Java 6, Eclipse, PL/SQL, SQL connector, Sub Version.

Confidential

Python Developer

Responsibilities:

  • Participated in the complete SDLC process and used PHP to develop website functionality.
  • Coding in LAMP (Linux, Apache, MySQL, and PHP) environment.
  • Developed GUI HTML, XHTML, AJAX, CSS and JavaScript(JQuery).
  • Built application logic usingPython, used the Django Framework to develop the application.
  • Used Django APIs for database access.
  • Description Bluetooth enabled camcorder Embedded in a headset to enable hands-free audio and video recording with optical zoom capability.
  • Rewrite existing Java application inPython module to deliver certain format of data
  • WrotePython scripts to parse XML documents and load the data in database.
  • Utilized PyQt to provide GUI for the user to create, modify and view reports based on client data.
  • UsedPython based GUI components for the front end functionality such as selection criteria.
  • Developed monitoring and notification tools usingPython.
  • Participated in requirement gathering and worked closely with the architect in designing and modeling.
  • Created Data tables to display customer information and add, delete, update customer records usingPython, MySQL and XHTML.
  • Used PyQt for the functionality filtering of columns helping customers to effectively view their transactions and statements. Implemented navigation rules for the application and page outcomes, written controllers using annotations.
  • Written queries in HQL and Native SQL and criteria API.
  • Added the navigations and paginations and filtering columns and adding and removing the desired columns for view utilizingPython based GUI components.
  • Implemented marshalling and UN marshalling XML to HTML and HTML to XML.
  • Created PyUnit test scripts and used for unit testing.
  • Actively participated in System Testing, production support and maintenance/patch deployments.
  • Worked on RUP development environment and used Rational ClearCase for versioning.
  • Used JQuery for selecting particular DOM elements when parsing HTML.
  • Developed SQL Queries, Stored Procedures, and Triggers Using Oracle 9i SQL, PL/SQL.
  • Developed test cases usingPython unit test, pylint and nose.

Environment: Python, HTML, JavaScript, Ajax, PyQT, PyUnit, PL/SQL, and Oracle SQLDeveloper

Confidential

Python Developer

Responsibilities:

  • Worked on requirement gathering and High level design.
  • Used HTML/CSS and Javascript for UI development.
  • Converted Visual basic Application toPython and MSQL.
  • UsedPython scripts to update content in the database and manipulate files.
  • Written many programs to parse excel file and process many user data with data validations.
  • Used Thales theorem for applying encryption and decryption of ISO standard message inPython programming.
  • Ensured high quality data collection and maintaining the integrity of the data.
  • Contributed patches back to Django.
  • UtilizedPython in the handling of all hits on Django, Redis, and other applications.
  • Developed object-oriented programming to enhance company product management.
  • Used severalPython libraries like wxPython, numPY and matPlotLib.
  • Was involved in environment code installation as well as the SVN implementation.
  • Build all database mapping classes using Django models.
  • Created unit test/regression test framework for working/new code.
  • Responsible for debugging and troubleshooting the web application.

Environment: Python2.6, Scipy, Pandas, Bugzilla, SVN, C++, Java, JQuery, MS SQL, Visual Basic, Linux, Eclipse, Java Script, XML, JASPER, PL/SQL, Oracle 9i, Shell Scripting, HTML5/CSS, Apache.

We'd love your feedback!