We provide IT Staff Augmentation Services!

Hadoop Developer Resume

0/5 (Submit Your Rating)

Conway, AR

SUMMARY

  • 7+ years of professional IT experience in analyzing requirements, designing, building, highly distributed mission critical products and applications.
  • 4+ years of data analytics experience in Apache Hadoop Cloudera and Hortonworks Distributions.
  • Solid experience as a Java /J2EE Developer developing business components using core Java concepts and classes like inheritance, polymorphism, collections, serialization, multithreading. etc.
  • Expertise in core Hadoop and Hadoop technology stack which includes HDFS, Map Reduce, Oozie, Hive, Sqoop, Pig, Flume, HBase, YARN, Spark, Storm, Kafka and Zookeeper.
  • Experience in AWS cloud environment and on S3 storage and EC2 instances and deploying in it.
  • Having Knowledge to implement Hortonworks (HDP 2.1, HDP 2.2 and HDP 2.3), Cloudera (CDH2, CDH3, CDH4 and CDH5) on Linux.
  • Well versed in installation, configuration, supporting and managing of Big Data and underlying infrastructure of Hadoop Cluster.
  • Experienced in implementing complex algorithms on semi/unstructured data using MapReduce programs.
  • Developed and implemented core API services using Scala and Spark.
  • Experienced in working with structured data using Hive QL, join operations, Hive UDFs, Partitions, bucketing and internal/external tables.
  • Experienced in migrating ETL kind of operations using Pig transformations, operations and UDFs.
  • Hands - on experience in Microsoft Azure Cloud Services (PaaS & IaaS), Storage, Web Apps, Active Directory, Application Insights, DocumentDB, Internet of Things (IoT).
  • 3+ experience in Data modelling, Data Migration, Data analysis and Data Mapping from source to target systems.
  • Experience in Data Ingestion, Processing, Development from Various RDBMS data sources into a Hadoop Cluster using Map Reduce/Pig/Hive/Sqoop
  • Wrote Python scripts to parse XML documents and load the data in database.
  • Configured different topologies for Stormcluster and deployed them on regular basis.
  • Experienced in implementing unified data platform to get data from different data sources using Apache Kafka brokers, cluster, Java producers and Consumers.
  • Excellent working knowledge of Spark Core, Spark SQL, Spark Streaming.
  • Hands on installing and configuring nodes CDH4 Hadoop Cluster on CentOS.
  • Experienced in working with in-memory processing frame work like Spark transformations, Spark and Spark streaming.
  • Ability to implement and deploy Azure offerings including both the IaaS and PaaS offering.
  • Experienced in proving User-based recommendations by implementing collaborative filtering and matrix factorization and different classification techniques like random forest, SVM, K-N Nusing Spark Mlib library.
  • Very good experience with application servers like WebLogic, WebSphere and Tomcat.
  • Used Python Flask micro-framework for workflow and Cassandra database for managing the raw, transformed data.
  • Understanding and knowledge of NOSQL databases like HBase, Cassandra, Mongo DB, Teradata and on Data warehouse.
  • Installed and configured Cassandra and has knowledge about Cassandra architecture, read, write paths and query.
  • Worked on AWS more TEMPthan 1 year to create EC2 instance and installed Java, Zookeeper and Kafka on those instances.Implemented cloud formation templates to create EC2 instances instead of creating manually.Worked on S3 buckets on AWS to store Cloud Formation Templates.
  • Involved in NoSQL (Datastax Cassandra) database design, integration and implementation and writing scripts and invoking them using CQLSH.
  • Involved in data modeling in Cassandra and involved in implementing sharding and replication strategies in MongoDB.
  • Developed fan-out workflow using Flume for ingesting data from various data sources like Web servers, RESTAPI by using different sources and ingesting data into Hadoop with HDFS sink.
  • Configured and deployed the application on the WebSphere server.
  • Experienced in implementing custom interceptors and sterilizers in Flume for specific customer requirements.
  • Implemented Azure Application Insights to store user activities and error logging.
  • Worked on importing and exporting data from MySQL into HDFS, HIVE and Hbase using Sqoop.
  • Monitored log input from several data centers, via Spark Stream, which was analyzed in Apache Storm and data was parsed and saved into Cassandra.
  • Developed Spark code using Scalaand Spark-SQL/Streaming for faster testing and processing of data. Strong Experience in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems MySQL and vice versa.
  • Understanding / knowledge of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and MapReduce programming paradigm.
  • Good exposure to Apache Hadoop MapReduce programming, Pig scripting and Distribute Application and HDFS.
  • Experience in managing Hadoop clusters using Cloudera Manager Tool.
  • Very good experience in complete project life cycle (design, development, testing and implementation) of Client Server and Web applications.
  • Experience in administering, installation, configuration, troubleshooting, security, backup, performance monitoring and fine tuning of Linux Redhat.
  • Worked on Cluster co-ordination services through Zookeeper.
  • Actively involved in coding using CoreJavaand collection APIs such as Lists, Sets and Maps.
  • Hands on experience in application development using Java, RDBMS, and Linux shell scripting.
  • Experience on Java Multi-Threading, Collection, Interfaces, Synchronization, and Exception Handling.
  • Involved in writing PL/SQL stored procedures, triggers and complex queries.
  • Worked in Agile environment with active scrum participation.

TECHNICAL SKILLS

Hadoop/Big Data: HDFS, MapReduce, Hbase, Pig, Hive, Sqoop, MongoDB, Cassandra, Flume, Oozie, Zookeeper, YARN, AWS, Spark, Falcon, Kafka, Teradata, Storm, Scala, ETL, Informatica.

Java & J2EE Technologies: Core Java, Servlets, JSP, JDBC, Java Beans, Maven, Gradle, JUnit, TestNG.

IDE’s: Eclipse, Net beans, Intellij Idea.

Frameworks: MVC, Struts, Hibernate and Spring.

Programming languages: C, C++, Java, Python, Ant scripts, Linux shell scripts

Databases: Oracle 11g/10g/9i, MYSQL, DB2, MS-SQL SERVER

Web Servers: WebLogic, WebSphere, Apache Tomcat

Web Technologies: HTML, XML, JavaScript, AJAX, SOAP, WSDL, JAX-RS, RESTful, JAX-WS.

Network Protocols: TCP/IP, UDP, HTTP, DNS, DHCP

Version Controls: CVS, SVN, GIT.

PROFESSIONAL EXPERIENCE

Confidential, Conway, AR

Hadoop Developer

Responsibilities:

  • Worked on analyzing Hadoop cluster and different big data analytic tools including Pig, Hbase database and Sqoop, Cassandra, Zookeeper, AWS.
  • Evaluated business requirements and prepared detailed specifications that followed project guidelines required to develop written programs.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it using MapReduce programs.
  • Implemented MapReduce programs to retrieve Top-K results from unstructured data set.
  • Migrated various Hive UDFs and queries into Spark SQL for faster requests as part of POC implementation.
  • Optimized Map Reduce Jobs to use HDFS efficiently by using various compression mechanisms.
  • Worked on reading multiple data formats on HDFS using Scala.
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and extracted the data from HDFS to MySQL using Sqoop.
  • Ability to create test automation for big data solution components, including data verification, data quality checks, code coverage analysis and data analysis.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Experience in AWS cloud environment and on S3 storage and EC2 instances
  • Developed fan-out workflow using Flume for ingesting data from various data sources like Webservers, RESTFul API by using different sources and ingesting data into Hadoop with HDFS sink.
  • Involved in migrating MongoDB version 2.4 to 2.6 and implementing new security features and designing more efficient groups.
  • Worked on implementing Hadoop Streaming, Python MapReduce for analytics.
  • Worked on Large node clusters and Strong Experience in Multi-node setup of Hadoop cluster.Analyzed Large amounts of data sets to determine optimal way to aggregate and report on it.
  • Developed Spark scripts by using Scala Shell commands as per the requirement.
  • Good experience in implementing Kerberos & Ranger in Hadoop Ecosystem.
  • Hands on installing and configuring nodes CDH4 Hadoop Cluster on CentOS.
  • Implemented various ETLsolutions as per the business requirement using Informatica
  • Experience creating ETL jobs to load JSON data and server data into MongoDB and transformed MongoDB into the Data Warehouse.
  • Involved in data modeling in Cassandra and MongoDB and in choosing indexes and primary keys based on the client requirement.
  • Configured Spark streaming to receive real time data from the Kafka and stored the stream data to HDFS using Scale.
  • Used Spark for Parallel data processing and better performance.
  • Responsible for building data solutions in Hadoop using Cascading frameworks.
  • Used Pig for data cleansing and extracting the data from the web server output files to load into HDFS.
  • Developed a data pipeline using Kafkaand Storm to store data into HDFS.
  • Implemented Kafka Java producers, create custom partitions, configured brokers and implemented High level consumers to implement data platform.
  • Implemented Storm topologies to preprocess data, implemented custom grouping to configure partitions.
  • Managed and reviewed Hadoop log files.
  • Involved in creating Hive tables, loading with data and writing Hive queries which will run internally in MapReduce way.
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
  • Installed and configured Pig and also wrote Pig Latin scripts.
  • Responsible for managing data coming from different sources.
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.

Environment: Hadoop, MapReduce, HDFS, Hive, Pig, Java, SQL Sqoop, Java (jdk 1.6), Spark, Kafka, AWS, MongoDB, Storm, Cassandra, Scala, ETL, Informatica, CentOS, Talend.

Confidential, Chicago, IL

Hadoop Developer

Responsibilities:

  • Installed and configured Cassandra and querying using Cassandra shell.
  • Worked on writing MapReduce jobs to discover trends in data usage by customers.
  • Worked on and designed Big Data analytics platform for processing customer interface preferences and comments using Java, Hadoop, Hive and Pig.
  • Involved in Hive-Hbase integration by creating Hive external tables and specifying storage as Hbase format.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Experienced in rewriting existing Python modules to deliver certain format of data.
  • Experienced in defining job flows to run multiple MapReduce and Pig jobs using Oozie.
  • Installed and configured Hive and also wrote Hive QL scripts.
  • Loaded the data into relational database for reporting, dash boarding and ad-hoc analyses, which revealed ways to lower operating costs and offset the rising cost of programming.
  • Created ETL jobs to load JSON data and server data into MongoDB and transformed MongoDB into the Data Warehouse.
  • Involved in ETL code deployment, performance tuning of mappings in Informatica, Talend.
  • Created reports and dashboards using structured and unstructured data.
  • Worked on Apache Ranger for HDFS, HBase, Hive access and permissions to the users through active directory.
  • Implemented HBase co-processors, Observers to work as event based analysis.
  • Hands on installing and configuring nodes CDH4 Hadoop Cluster on CentOS.
  • Implemented Hive Generic UDFs to implement business logic.
  • Accessed Hive tables to perform analytics from Java applications using JDBC.
  • Ran batch processes using Pig Scripts and developed Pig UDFs for data manipulation according to Business Requirements.
  • Experience with streaming work flow operations and Hadoop jobs using Oozie workflow and scheduled through AUTOSYS on a regular basis.
  • Worked on implementing Hadoop Streaming, Python MapReduce for analytics.
  • Developed Spark SQL scripts and involved in converting Hive UDFs to Spark SQL UDFs.
  • Responsible for batch processing and real time processing in HDFS and NOSQL Databases.
  • Responsible for retrieval of data from Casandra and ingestion to Pig.
  • Experience in customizing MapReduce framework at various levels by generating Custom Input formats, Record Readers, Partitioners and Data types.
  • Experienced with multiple files in HIVE, AVRO, Sequence file formats.
  • Created and maintained Technical documentation for launching HADOOP Clusters and for executing Pig Script.
  • Implemented business logic by writing Pig UDFs in Java and used various UDFs from Piggybanks and other sources.

Environment: Casandra, Map jobs, Spark SQL, ETL, Pig Scripts, Flume, Hadoop BI, Pig UDFs, Oozie, AVRO, Hive, Map Reduce, Scala, Java, Eclipse, Zookeeper, CentOS and Informatica.

Confidential, Birmingham, AL

Hadoop Developer

Responsibilities:

  • Involved in the complete Software Development Life Cycle (SDLC) to develop the application.
  • Analyzed large node cluster and used different big data analytic tools including Pig, Hbase database and Sqoop, Cassandra and Zookeeper.
  • Involved in loading data from LINUX file system to HDFS.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Imported and exported data into HDFS and Hive using Sqoop.
  • Implemented test scripts to support test driven development and continuous integration.
  • Developed multiple MapReduce jobs in Java for data cleaning.
  • Installed and configured Hadoop Map Reduce, HDFS; developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
  • Created Pig Latin scripts to sort, group, join and filter the enterprise wide data.
  • Created Hive tables, loaded with data and wrote Hive queries that will run internally in MapReduce way.
  • Supported MapReduce programs that are running on the cluster.
  • Analyzed large data sets by running Hive queries and Pig scripts.
  • Worked on tuning the performance Pig queries.
  • Mentored analyst and test team on writing Hive queries.
  • Installed Oozie workflow engine to run multiple MapReduce jobs.
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
  • Worked on Zookeeper for coordinating between different master node and data nodes.

Environment: Hadoop, HDFS, Map Reduce, Hive, Pig, Sqoop, Linux, Java, Oozie, Hbase, Zookeeper.

Confidential, Raleigh, NC

Java /J2EE Developer

Responsibilities:

  • Worked with business users to determine requirements and technical solutions.
  • Followed agile methodology (Scrum Standups, Sprint Planning, Sprint Review, Sprint Showcase and Sprint Retrospective meetings).
  • Developed business components using core Java concepts and classes like Inheritance, Polymorphism, Collections, Serialization, Multithreading, etc.
  • Used spring framework that handles application logic and makes calls to business make them as Spring Beans.
  • Implemented, configured data sources, session factory and used Hibernate template to integrate Spring with Hibernate.
  • Developed Web Services to allow communication between applications through SOAP over HTTP with JMS and Mule ESB.
  • Actively involved in coding using CoreJavaand collection APIs such as Lists, Sets and Maps
  • Developed a Web Service (SOAP, WSDL) that is shared between front end and cable bill review system.
  • Implemented Rest based Web Service using JAX-RS annotations, Jersey implementation for data retrieval with JSON.
  • Developed Maven scripts to build and deploy the application onto WebLogic application Server and ran UNIX shell scripts and implemented autodeployment process.
  • Used Maven as the build tool and is scheduled/triggered by Jenkins (build tool).
  • Develop JUNIT test cases for application unit testing.
  • Implement Hibernate for data persistence and management.
  • Used SOAP UI tool for testing Web Services connectivity.
  • Used SVN as version control to check in the code; created branches and tagged the code in SVN.
  • Used RESTful Services to interact with the Client by providing the RESTFul URL mapping.
  • Used Log4j framework to log/track application and debugging.

Environment: JDK 1.6, Eclipse IDE, Core Java, J2EE, Spring, Hibernate, Unix, Web Services, SOAP UI, Maven, Web logic Application Server, SQL Developer, Camel, Junit, SVN, Agile, SONAR, Log4j, REST, Log 4j, Teradata, JBPM.

Confidential

Java Developer

Responsibilities:

  • Involved in analysis, design and development of Expense Processing system.
  • Created used interfaces using JSP.
  • Developed the Web Interface using Servlets, Java Server Pages, HTML and CSS.
  • Developed the DAO objects using JDBC.
  • Business Services using the Servlets and Java.
  • Design and development of User Interfaces and menus using HTML 5, JSP, Java Script, client side and server side validations.
  • Developed GUI using JSP, Struts framework.
  • Involved in developing the presentation layer using Spring MVC/Angular JS/JQuery.
  • Involved in designing the user interfaces using Struts Tiles Framework.
  • Used Spring 2.0 Framework for Dependency injection and integrated with the Struts Framework and Hibernate.
  • Used Hibernate 3.0 in data access layer to access and update information in the database.
  • Experience in SOA (Service Oriented Architecture) by creating the Web Services with SOAP and WSDL.
  • Developed JUnit test cases for all the developed modules.
  • Used Log4J to capture the log that includes runtime exceptions, monitored error logs and fixed the problems.
  • Used RESTful Services to interact with the Client by providing the RESTful URL mapping.
  • Used CVS for version control across common source code used by developers.
  • Used ANT scripts to build the application and deployed on WebLogic Application Server 10.0.

Environment: Struts1.2, Hibernate3.0, Spring2.5, JSP, Servlets, XML, SOAP, WSDL, JDBC, JavaScript, HTML, CVS, Log4J, JUNIT, WebLogic App server, Eclipse, Oracle, Restful.

We'd love your feedback!