We provide IT Staff Augmentation Services!

Hadoop Developer Resume

3.00/5 (Submit Your Rating)

New York, NY

PROFESSIONAL SUMMARY:

  • 10 years of IT experience in software Development and Big Data Technologies and Analytical Solutions with 2+ years of hands - on experience in development and design of Java and related frameworks and 1+ years’ experience in design, architecture, and data modeling as database developer.
  • Over 4 years’ experience as Hadoop Developer with good knowledge of Hadoop framework, Hadoop Distributed file system and Parallel processing implementation.
  • Built and Deployed Industrial scale Data Lake on on premise and Cloud platforms .
  • Experienced in Hadoop Ecosystems HDFS, Map Reduce, Hive, Pig, Python, HBase, Couchbase , Sqoop, Hue, Oozie, Impala, Spark .
  • Excellent understanding / knowledge of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
  • Experienced in handling different file formats like Text file, Avro data files, Sequence files, Xml and Json files.
  • Experience with Testing Map Reduce programs using MRUnit, Junit, ANT, Maven and EasyMock.
  • Expertise in deployment of Hadoop, Yarn, Spark integration with Cassandra, Ignite and RabbitMQ, Kafka etc.
  • Designed and developed Enterprise-level data, integration, and reporting solutions.
  • Good knowledge in using apache NiFi to automate the data movement between different Hadoop systems.
  • Upgraded Hadoop CDH to 5.x, and Hortonworks.
  • Installed, Upgraded and Maintained Cloudera Hadoop-based software.
  • Experience with hardening Cloudera Clusters, Cloudera Navigator and Cloudera Search.
  • Deployed Search engine for Big Data.
  • Processed query and transaction, optimized query on Big Data.
  • Ingested data and performed update/insert/delete records in Cloudera.
  • Experience in managing and reviewing Hadoop log files.
  • Hands on experience in Import/Export of data using Hadoop Data Management tool SQOOP.
  • Experience in importing streaming logs and aggregating the data to HDFS through Flume .
  • Strong experience in writing Map Reduce programs for Data Analysis. Hands on experience in writing custom practitioners for Map Reduce.
  • Analyzed data with  Hive, Pig  and scrubbed HBase  Data and processed with  Oozie .
  • Involved in moving all log files generated from various sources to HDFS and Spark for further processing.
  • Excellent understanding and knowledge of NOSQL databases like MongoDB, HBase, Couchbase and Cassandra.
  • Designed and deployed AWS solutions using E2C, S3, RDS, EBS, Elastic Load Balancer, Auto scaling groups, Opsworks
  • Experience in implementing Kerberos authentication protocol in Hadoop for data security.
  • Experience with distributed systems, large-scale non-relational data stores, RDBMS, NoSQL map-reduce systems, data modeling, database performance, and multi-terabyte data warehouses.
  • Experience in Dimensional modelling, logical modelling and Physical data modelling .
  • Experience in Software Development Life Cycle (Requirements Analysis, Design, Development, Testing, Deployment and Support).
  • Experienced with REST APIs based on frameworks such as Scalatra, CXF, Sinatra, Spray.
  • Experienced with code versioning and dependency management systems such as Git, SVT, and Maven.
  • Experienced with AWS like EC2, S3, EMR, OpenStack cloud infrastructures.
  • Experienced in working with scheduling tools such as UC4, Cisco Tidal enterprise scheduler, or Autosys.

TECHNICAL SKILLS: 

Hadoop ECO Systems: Hadoop, MapReduce, HDFS, HBase, Hive, Pig, Sqoop, ZooKeeper, Flume, Impala, Hue, Oozie, Cloudera Manager, Accumulo, Spark, Kafka, Akka and MRUnit.

Analytics Softwares: R, SAS, Matlab

NO SQL: MongoDB, Couchbase, Cassandra

Data Bases: MS SQL Server 2000/2005/2008/2012 , MY SQL, Oracle 9i/10g, MS access, Teradata TeradataV2R5

Languages: Languages Java JDK1.4 1.5 1.6 (JDK 5 JDK 6), C/C++, SQL, Teradata SQL, PL/SQL.

Operating Systems: Windows Server 2000/2003/2008 , Windows XP/Vista, Mac OS, UNIX, LINUX

Java Technologies: Servlets, JavaBeans, JDBC, JNDI, JTA, JPA, E

Frame Works: Jakarta Struts 1.1, JUnit and JTest, LDAP,   Scalatra, CXF, Sinatra, Spray

IDE’s & Utilities: Eclipse, Maven, NetBeans.

SQL Server Tools: SQL Server Management Studio, Enterprise Manager, Query Analyser, Profiler, Export & Import (DTS).

Web Technologies: ASP.NET, HTML,XML

Testing & Case Tools: Bugzilla, QuickTestPro(QTP)9.2, Selenium, Quality Center, Test Link, Junit, Log4j, Rational Clear case, ANT.

Business Intelligence Tools: Tableau, Pentaho, Qlikview, Micro Strategy, Business Objects

ETL Tools: Informatica, Infosphere, TalenD

Methodologies: Agile, UML, Design Patterns

PROFESSIONAL EXPERIENCE:

Confidential, New York, NY

Hadoop Developer

Responsibilities:

  • Invovled in Connected Innovation and Big Data Analytics enterprise priority for global BP.
  • Experience with Agile development processes and practices.
  • Working in agile, successfully completed stories related to ingestion, transformation and publication of data on time.
  • Designed/Created HDFS data lake by drawing relationship between different sources of data from various systems
  • Developed Data Lake architecture capable of consuming, processing and storing logs from sensors of devices.
  • Establishment of Data Lake using Hadoop platform is aiming to create an Enterprise grade ecosystem that enables Analytics and Global Reporting resulting into data driven Decision Making, Flexible Management, Financial Reporting, Data Governance across all business units.
  • Expertise in designing and deployment of Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, Oozie, ZooKeeper, Sqoop, flume, Apache Spark, Impala with Hortonworks Distribution.
  • Involved in loading and transforming large sets of structured, semi-structured and Unstructured data and analyzed them by running Hive queries and Pig scripts.
  • Created Analytics and reports from data using the HiveQL
  • Implemented various data Importing and exporting jobs into HDFS and Hive using Sqoop.
  • Started using apache NiFi to copy the data from local file system to HDFS.
  • Transform and created RDDs, DataFrame using Spark.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, PySpark and Scala.
  • Analyzed the SQL scripts and designed the solution to implement using PySpark
  • Developed Spark scripts by using Scala shell commands as per the requirement.
  • Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
  • Developed Scala scripts, UDFFs using PyS park, Data frames/SQL and RDD/MapReduce in Spark 1.3 for Data Aggregation, queries and writing data back into OLTP system directly or through Sqoop.
  • Exploring with Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark context, spark-SQL, Data Frame, pair RDD's, Spark YARN.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
  • Extensively used scripting (python and shell) to provision and spin up virtualized Hadoop clusters.
  • Implemented Fair schedulers on the Job tracker with appropriate parameters to share the resources of the Cluster for the Map Reduce jobs given by the users.
  • Involved in creating Hive tables, loading the data using it and in writing Hive queries to analyze the data.
  • Worked on tuning the performance Pig queries.
  • Involved in loading data from LINUX file system to HDFS.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Imported streaming logs and aggregating the data to HDFS through Flume.
  • Experience working on processing unstructured data using Pig and Hive.
  • Extensively used Pig for data cleansing, data deduplication .
  • Created partitioned tables in Hive.
  • Managed and reviewed Hadoop log files.
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in MapReduce way.
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
  • Developed bash scripts to bring the Tlog files from ftp server and then processing it to load into hive tables.
  • S cheduled bash scripts using Resource Manager Scheduler.
  • Installed and configured Pig and also written Pig Latin scripts.
  • Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
  • Created UDFs to calculate the pending payment for the given Residential or Small Business customer, and used in Pig and Hive Scripts.
  • Developed multiple Map Reduce jobs in java for data cleaning and preprocessing .

Environment: Cloudera, Hive, MapReduce, Agile, Sqoop, NiFi, Flume, Oozie, Spark, Pig, Scala, Linux, Java, Python, PySpark, bash, UNIX Shell Scripting and Big Data

Confidential, Holmdel, NJ

Hadoop Developer

Responsibilities:

  • Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.
  • Worked within and across Agile teams to design, develop, test and support technical solutions across a full-stack of development tools and technologies.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Developed Simple to complex Map reduce Jobs using Hive and Pig.
  • Optimized Map Reduce Jobs to use HDFS efficiently by using various compression mechanisms
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and Extracted the data from MySQL into HDFS using Sqoop.
  • Wrote data ingestion systems to pull data from traditional RDBMS platforms such as Oracle and Teradata and store it in NoSQL databases such as MongoDB, Couchbase , Cassandra.
  • Experience with Cassandra, with ability to drive the evaluation and potential implementation of it as a new platform.
  • Implemented analytical engines that pull data from API data sources and then present data back as either an API or persist it back into a NoSQL platform.
  • Involved in moving all log files generated from various sources to HDFS and Spark for further processing.
  • Involved in requirement and design phase to implement Streaming Lambda Architecture to use real time streaming using Spark and Kafka.
  • Developed real-time data synchronization systems with reactive programming concepts like Akka and Kafka  .
  • Implemented and released of the first version of the IOT cloud platform, using Apache Storm, Kafka, MQTT, HBase, MongoDB and Java EE.
  • Integrated Apache Storm with Kafka to perform web analytics. Uploaded click stream data from Kafka to Hdfs, Hbase and Hive by integrating with Storm.
  • Created User defined types to store specialized data structures in Couchbase .
  • Implemented a distributed messaging queue to integrate with Cassandra using Apache Kafka and ZooKeeper
  • Experienced in using Avro data serialization system to handle Avro data files in map reduce programs.
  • Design, implementation, test, debug of ETL mappings and workflows.
  • Develop ETL routines to source data from client source systems and target the data warehouse.
  • Configure ETL tool and ensures Full and Incremental loads run successfully and be familiar with 'change data capture' concepts.
  • Experience with Python and Shell scripting, which will be used for automating ETL jobs and tasks.
  • The data is collected from distributed sources into Avro models . Applied transformations and standardizations and loaded into Hive for further data processing.
  • Built Platfora Hadoop multi-node cluster test labs using Hadoop Distros (CDH 4/5, Apache Hadoop, MapR
  • and HortonWorks) and Hadoop Eco-systems, Virtualizations and Amazon Web Services component.
  • Installed, Upgraded and Maintained Cloudera Hadoop-based software.
  • Experience with hardening Cloudera Clusters, Cloudera Navigator and Cloudera Search.
  • Managing Running Jobs, Scheduling Hadoop Jobs, Configuring the Fair Scheduler, Impala Query Scheduling.
  • Deployed Search engine for Big Data.
  • Developed the XML Schema and Web services for the data maintenance and structures.
  • Implemented the Web Service client for the login authentication, credit reports and applicant information using Apache Axis 2 Web Service.
  • Developed workflows using custom MapReduce, Pig, Hive and Sqoop.
  • Built reusable Hive UDF libraries for business requirements which enabled users to use these UDF's in Hive Querying.
  • Performed troubleshooting, fixed and deployed many Python bug fixes of the two main applications that were a main source of data for both customers and internal customer service team.
  • Comprehensive knowledge in understanding different components of Spark framework
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Participated in development/implementation of  Cloudera Hadoop  environment.
  • Integrated Hadoop Security with Active Directory by implementing Kerberos for authentication and Sentry for authorization.
  • Used struts validation framework for form level validation
  • Wrote test cases in Junit for unit testing of classes.

Environment: Hadoop, HDFS, Pig, Agile, Cloudera, Accumulo, Cassandra, MongoDB, Couchbase, Sqoop, Scala, Python, Spark, MQTT, Storm, Kafka, Akka, Kerberos, JSP, HTML, XML, ANT 1.6, Perl, Python, JavaScript, Junit 3.8, Avro, Impala, Hue.

Confidential, Jacksonville, Florida

Hadoop Developer

Responsibilities:

  • Installed and configured Hadoop and Hadoop stack on a 16 node cluster.
  • Worked on analyzing Hadoop cluster using different big data analytic tools including Pig, Hive, and Map Reduce.
  • Worked on debugging, performance tuning of Hive & Pig Jobs.
  • Developed data access libraries that bring MapReduce, Graph, and RDBMS data to users of Scala, Java, Python.
  • Analyze large and critical datasets using Cloudera, HDFS, Hbase, MapReduce, Hive, Hive UDF, Pig, Sqoop, Zookeeper, & Mahout.
  • Designed and deployed AWS solutions using E2C, S3, RDS, EBS, Elastic Load Balancer, Auto scaling groups, Opsworks
  • Involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs.
  • Implemented Cluster Coordination services through Zookeeper.
  • Worked on Cisco Tidal Enterprise Scheduler (TES) which is a friendlier alternative to Oozie, the native Hadoop scheduler. Built-in Cisco TES connectors to Hadoop components eliminate manual steps such as writing Sqoop code to download data to HDFS and executing a command to load data to Hive.
  • Install, configure, and operate data integration and analytic tools i.e. Informatica, Chorus, SQLFire, & Gem Fire XD for business needs.
  • Worked with file formats TEXT, AVRO, PARQUET and SEQUENCE files .
  • Develop scripts to automate routine DBA tasks (i.e. refresh, backups, vacuuming, etc.)
  • Installed and configured Hive and also wrote Hive yarn’s that helped spot market trends.
  • Used Hadoop streaming to process terabytes data in XML format.
  • Involved in loading data from UNIX file system to HDFS.
  • Design, develop, unit test, and support ETL mappings and scripts for data marts using Talend.
  • Dynamic schema validation & custom data cleansing within Talend processes.
  • Design and development of Talend jobs to load metadata in SQL Server data base.
  • Developed complex Talend ETL job to load the data from file to HDFS, HDFS to Hive, Hive to Oracle and Oracle to DataMart.
  • Experienced with REST APIs based on frameworks such as Scalatra, CXF, Sinatra, Spray.
  • Involved in Developing a Restful API'S service using Python Flask framework.
  • Experienced with code versioning and dependency management systems such as Git, SVT, and Maven.
  • Experienced with Hive customization, i.e. UDFs, UDTFs and UDAFs.
  • Experienced with Python-Hive integration including Pandas, Numpy and Scipy.
  • Experienced with AWS like EC2, S3, EMR, OpenStack cloud infrastructures.
  • Experienced in working with scheduling tools such as UC4, Cisco Tidal enterprise scheduler, or Autosys.
  • Knowledge of BI tools like Pentaho and Tableau.

Environment: CDH4 with Hadoop 1.x, HDFS, Pig, Cloudera, Hive, Hbase, zookeeper, MapReduce, Java, Sqoop, Oozie, Linux, UNIX Shell Scripting and Big Data, Python, Flask, Cisco Tidal enterprise scheduler, OpenStack, Pentaho, Tableau, TalenD

Confidential, Columbus, OH

Java/J2ee Developer

Responsibilities:

  • Understanding and analyzing the project requirements.
  • Analysis and Design with UML and Rational Rose.
  • Created Class Diagrams, Sequence diagrams and Collaboration Diagrams
  • Used the MVC architecture.
  • Worked on Jakarta Struts open framework.
  • Wrote spring configuration for the beans defined and properties to be injected into themusing spring's Dependency Injection.
  • Implemented a ftp utitlity program for copying the contents of an entire directory recursively upto two levels from a remote location using Socket Programming.
  • Implemented a reliable socket interface using the sliding window protocol like TCP stream sockets over UDP unreliable communication channel and later on, tested using the Ftp utility program.
  • Strong domain knowledge of TCP/IP with the expertise in socket programming and IP security domain (IPSec, TLS, SSl and VPN, Firewall and NATs).
  • Have built the strong communication between the source and destination message using socket programming.
  • Hands on experience in writing Spring Restful Web services using JSON / XML.
  • Developed the Spring Features like Spring MVC, Spring DAO, Spring Boot, Spring Batch, Spring Security.
  • Using AngularJS, HTML5, CSS3 all HTML and DHTML is accomplished through AngularJS directives.
  • Developed Servlets in order to deal with requests for account activity,
  • Developed Controller Servlets and Action Servlets to handle the requests and responses.
  • Developed Servlets and created JSP pages for viewing on a HTML page.
  • Developed the front end using JSP.
  • Developed various EJB's to handle business logic.
  • Designed and developed numerous Session Beans deployed on Web logic Application Server.
  • Implemented Database interactions using JDBC with back-end Oracle.
  • Worked on Database designing, Stored Procedures, and PL/SQL.
  • Created triggers and stored procedures using PL/SQL.
  • Written queries to get the data from the Oracle database using SQL.

Environment: J2EE, Servlets, JSP, Struts, Spring Restful WebServices, NATS, Hibernate, Oracle, TOAD, Web logic Server, AngularJS, HTML5, CSS3 all HTML and DHTML

Confidential, Boston, MA

SQL/JAVA Developer

Responsibilities:

  • Involved in complete requirement analysis, design, coding and testing phases of the project.
  • Implemented the project according to the Software Development Life Cycle (SDLC).
  • Developed JavaScript behavior code for user interaction.
  • Used HTML, JavaScript, and JSP and developed UI.
  • Used JDBC and managed connectivity, for inserting/querying& data management including stored procedures and triggers.
  • Designed the logical and physical data model, generated DDL scripts, and wrote DML scripts for Sql Server database.
  • Implemented application using JSP, Spring MVC, Spring IOC, Spring Annotations, Spring AOP, Spring Transactions, Hibernate.
  • Transformed project data requirements into project data models using Erwin.
  • Involved in logical and physical designs and transforms logical models into physical implementations.
  • Enhanced existing data model based on the requirements, maintained data models in Erwin Model Manager.
  • Part of a team which is responsible for metadata maintenance and synchronization of data from database.
  • Involved in the design and coding of the data capture templates, presentation and component templates.
  • Designed and developed various SSIS packages (ETL) to extract and transform data and involved in Scheduling SSIS Packages.
  • Designed SSIS Packages to extract, transfer, load (ETL) existing data into SQL Server from different environments for the SSAS cubes.
  • Created reports using SQL Reporting Services (SSRS) for customized and ad-hoc Queries.
  • Extensively designed the packages and data mapping using Control flow task, Sequence container task, Dataflow Task, Execute SQL Task, Data conversion task, Derived Column task and Script Task in SSIS Designer.
  • Developed SSIS packages to automate the nightly extract transform and load jobs.
  • Worked with different methods of logging in SSIS.
  • Experience in creating complex SSIS packages using proper control and data flow elements.
  • Generated SSRS Report through SSIS Package using script component as per business requirement.
  • Wrote production implementation documents and provided Production support when required. Implemented MS SQL Server Analysis Services setup, tuning, cube partitioning, dimension design including hierarchical and slowly changing dimensions.
  • Designed STAR SCHEMA following Dimensional Modeling approach.
  • Designed Dimensional model to support Business process to answer complex business questions
  • Provided assistance to development teams on Tuning Data, Indexes and Queries.
  • Developed an API to write XML documents from database.
  • Used JavaScript and designed user-interface and checking validations.
  • Developed JUnit test cases and validated users input using regular expressions in JavaScript as well as in the server side.
  • Developed complex SQL stored procedures, functions and triggers.
  • Mapped business objects to database using Hibernate.
  • Wrote SQL queries, stored procedures and database triggers as required on the database objects.

Environment: Java, spring, XML, Hibernate, SQL Server, Maven2, JUnit, SSIS, SSAS, SSRS, Dimensional modelling, logical modelling and Physical data modelling

Confidential

SQL Server Developer

Responsibilities:

  • Actively involved in different stages of Project Life Cycle.
  • Documented the data flow and the relationships between various entities.
  • Actively participated in gathering of User Requirement and System Specification.
  • Created new Database logical and Physical Design to fit the new business requirement and implemented the same using SQL Server.
  • Created Clustered and Non-Clustered Indexes for improved performance.
  • Created Tables, Views and Indexes on the Database, Roles and maintained Database Users.
  • Developed new Stored Procedures, Functions, and Triggers.
  • Implemented Backup and Recovery of the databases.
  • Actively participated in User Acceptance Testing, and Debugging of the system.

Environment: Windows 2000 adv. server, Windows 2000/XP, MS SQL Server 2000, IIS, MS Visual Studio.

We'd love your feedback!