Hadoop/java Developer Resume
Milwaukee, WI
SUMMARY:
- Having 9 plus years of professional IT experience which includes 5 plus years of experience in Hadoop Development and Administration usingHortonworks And Cloudera Distributions.
- Over five plus years of experience in design, development, maintenance and support of Big Data using Hadoop(Cloudera and Hortonworks) Ecosystem tools like HDFS, Hive, Pig, Sqoop, Flume, Zookeeper, MapReduce, Spark, Oozie and Talend.
- Experienced in installation, configuration, supporting and monitoring Hadoop cluster using Cloudera manager and Hortonworks Ambari distributions.
- Have experience in installing, configuring, performance tuning and administrating Hadoop cluster for major Hadoop distributions like HDP 2.4.
- Strong working experience with ingestion, storage, querying, processing and analysis of big data.
- Experience in installation, configuration, supporting and managing Hadoop clusters.
- Expertise in writing Hadoop Jobs for analyzing data using Hive and Pig.
- Loaded streaming log data from various web servers into HDFS using Flume.
- Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems (RDBMS) and vice - versa.
- Knowledge of job workflow scheduling and monitoring tools like Oozie and Azkaban.
- Good understanding of NoSQL databases.
- Having experience on using HIVE, PIG.
- Load and transform large sets of structured, semi-structured and unstructured data using Hadoop ecosystem components.
- Experience in working with different data sources like Flat files, XML files And Databases.
- Worked with project documentation and also documented other application Related issues, bugs on internal wiki website.
- Full life cycle experience, involved in requirement analysis, design, Development, testing, deployment and support.
- Experience in analyzing logs for troubleshooting java application issues from Server side.
- A very good team player with the ability to work independently with minimal Supervision.
CORE COMPETENCIES:
- Application Development
- Object Oriented Programming (OOP)
- Big Data / Hadoop
- AWS
- Hadoop(HDP, Cloudera)
- HDFS
- Map Reduce, YARN
- Sqoop, Hive, Pig
- Flume, Impala
- Oozie, Zookeeper
- Spark, Tez, Ambari
- Talend
- PERL, PUPPET
- JBoss/WildFly,
- Websphere, Weblogic
- Apache Http
- Tomcat
TECHNICAL SKILLS:
- J2SE, J2EE
- XML Web Services SOAP/REST
- Eclipse, NetBeans,IntelliJ IDE
- HTML/CSS, Java Script
- Oracle 10g, 11i
- DB2, MySQL, MS Access
- SQL Server, Netezza
- Unix, Linux
- AIX, Solaris, Windows
- IP Center, JIRA
- Wily Introscope
- Nagios, UC4, Git
- Maven, ANT, Jenkin
- Subversion, Tortoise
- Azkaban, Filezilla
- Ambari
- Cloudera Manager
PROFESSIONAL EXPERIENCE:
Confidential, Milwaukee, WI
Hadoop/Java Developer
Responsibilities:
- Worked closely with Business Analysts to review the business specifications of the project and also to gather the ETL requirements.
- Working experience on designing and implementing complete end-to-end Hadoop Infrastructure including Pig, Hive, Sqoop, Oozie and Zookeeper.
- Used Talend to generate optimized code to load, transform, enrich, and cleanse data inside Hadoop.
- Created Talend jobs to copy the files from one server to another and utilized Talend FTP components.
- Created and managed Source to Target mapping documents for all Facts and Dimension tables.
- Involved in writing SQL Queries and used Joins to access data from Oracle, and MySQL.
- Prepared ETL mapping Documents for every mapping and Data Migration document for smooth transfer of project from development to testing environment and then to production environment.
- Create a table inside RDBMS, insert some data after load the same table into HDFS, Hive using Sqoop.
- Used Sqoop to import the data from RDBMS to Hadoop Distributed File System (HDFS) and later analyzed the imported data using Hadoop Components.
- Worked in designing parameters for extraction, cleansing, validation and transformation of data from various source systems to Data Warehouse.
- Used OOZIE Operational Services for batch processing and scheduling workflows dynamically.
- Involved in exporting the analyzed data to the databases such as Teradata, MySQL and Oracle using Sqoop for visualization and to generate reports for the BI team.
- Worked on Oozie scheduler to automate the pipeline workflow and orchestrate the sqoop, hive and pig jobs that extract the data on a timely manner.
- Developed customized Hive UDFs and UDAFs in Java, JDBC connectivity with hive development and execution of Pig scripts and Pig UDF's.
- Developed MapReduce jobs in java for data cleaning and preprocessing.
- Involved in creating Oozie workflow and Coordinator jobs for Hive jobs to kick off the jobs on time for data availability.
- Migrated HiveQL to SparkQL to validate Spark's performance with Hive's.
- Automated tasks using UNIX shell scripts.Responsible for developing, support and maintenance for the ETL (Extract, Transform and Load) processes using Talend Integration Suite.
- Used Talend Admin Console Job conductor to schedule ETL Jobs on daily, weekly, monthly and yearly basis.
- Worked Extensively on Talend Admin Console and Schedule Jobs in Job Conductor.
Environment: Hortonworks, Talend, Hadoop, HDFS, Pig, Hive, Sqoop, Shell Scripting, Core Java, RHEL.
Confidential, Foster City, CA
Hadoop Developer
Responsibilities:
- Working with the Data Frames and RDD's.
- Creating the tables in Hive and integrating data between Hive & Spark.
- Hands on experience with Spark Scala programming and good understanding of its 'In Memory' processing capability.
- Worked on creating the RDD's, DF's for the required input data and performed the data transformations using Spark Scala.
- Experienced with batch processing of data sources using Apache Spark.
- Experienced in working with RDDs.
- Monitoring and Debugging Hadoop jobs/Applications running in production.
- Worked on Providing User support and application support on Hadoop Infrastructure.
- Installed and configured Hadoop components Hive, Impala, Pig.
- Cluster maintenance as well as creation and removal of nodes.
- Monitor Hadoop cluster connectivity and security.
- Manage and review Hadoop log files.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Responsible for cluster maintenance, adding and removing cluster nodes, cluster monitoring and troubleshooting, manage and review data backups, manage and review Hadoop log files.
- Installed Oozie workflow engine to run multiple Hive and pig jobs.
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
- Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
- Diligently teaming with the infrastructure, network, database, application and business intelligence teams to guarantee high data quality and availability.
Environment: Cloudera Hadoop, HDFS, Pig, Hive, Sqoop, Shell Scripting, Java, UNIX.
Confidential, Atlanta, GA
Hadoop Developer
Responsibilities:
- Data Ingestion has done from the different sources like Oracle DB, Netezza and flat files.
- Written a master script to download files from FTP source,uncompress them into staging and do a HDFS put.
- Analyzing the data and using PIG, HIVE for the loading of the data into HDFS.
- Vast use of Shell scripting for the loading of data into HDFS.
- Developed customized flume agents to consume live network data and persist into HDFS.
- Designed hive partitions that get created on a daily basis using Oozie workflows.
- Designed PIG scripts to filter network data based on configuration settings in a MySQL database.
- Developed Oozie workflows to look for configuration changes and rewrite partitions for hive tables.
- Written a Java code for Schema Analysis Engine to analyze the data from different sources and create Hive QL scripts.
- Created Hive pre and partition tables for different sources manually and using Schema Analysis Engine.
- Created hive tables for various types of SERDE format.
- Worked on HCatalog which allows PIG and Map Reduce to take advantage of the SerDE data format transformation definitions that write for HIVE.
- Data Ingestion has done in three layers with different transformations using Hive and Pig Scripts.
- Written managed and external tables using Hive for different sources manually and using Schema Analysis Engine.
- Written Insert Hive QL scripts and Sqoop Jobs.
- Worked on different file formats (orc file,rcfile, sequence file, text file) and different Compression Codec’s(gzip,snappy).
- Worked on both Cloudera and Hortonworks distribution.
- Validated the data in Production and QA after upgrading the Clouderaversion of Hadoop Environment.
- Importing and exporting data into HDFS and Hive using Sqoop.
- HiveQL scripts to create, load, and query tables in a Hive.
- Installed and configured Hive and also written Hive UDFs.
- Created and executed Azkaban flow for different sources in DEV, QA and PROD.
- Worked on PIG Latin Scripts and UDF's while ingestion, querying, processing and analysis of Data.
- Validated daily volume metrics in DEV, QA and PROD for different sources.
- Experienced in managing and reviewing Hadoop log files.
- Responsible for naming conventions of column and table names in Hive and Pig Scripts.
- Load and transform large sets of structured, semi structured and unstructured data.
- Responsible to manage the data coming from different sources.
- Created and maintained Technical documentation for launching Hadoop clusters and for executing Hive queries and Pig Scripts.
- Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
Environment: Hortonworks/ClouderaHadoop, HDFS, Pig, Hive, Sqoop, Shell Scripting, Core Java, Netezza, Oracle 11g, Linux, UNIX.
Confidential, Chicago, IL
Big Data Developer
Responsibilities:
- Customized flume agents to consume live network data and persist into HDFS.
- Managing and reviewing Hadoop log files.
- Extracting files through Sqoop and place in HDFS and processed.
- Running Hadoop streaming jobs to process terabytes of xml format data.
- Loading and transforming large sets of structured, semi structured and unstructured data.
- Responsible to manage data coming from different sources.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Hive QL scripts to create, load, and query tables in a Hive.
- Installed and configured Hive and also written Hive UDFs.
- Designed hive partitions that get created on a daily basis using oozie workflows.
- Utilized Apache Hadoop environment by Cloudera works.
- Experienced in defining job flows
- Implemented automated system used to set infrastructure leveraged by different users of combine site in Java/Scala.
- Experienced in managing and reviewing Hadoop log files.
- Load and transform large sets of structured, semi structured and unstructured data
- Responsible to manage data coming from different sources.
- Supported Map Reduce Programs those are running on the cluster.
- Involved in loading data from UNIX file system to HDFS.
- Loading and transforming large sets of structured, semi structured and unstructured data.
- Developed Oozie workflows to look for configuration changes and rewrite partitions for hive tables.
- Worked on using technologies like Oracle Weblogic,RedhatJboss,WildFly, IBM Websphere administration for various clients.
Environment: Java, Hadoop, Hive, Pig, JDBC, UNIX, HTML, CSS, XML, Oracle Weblogic,RedhatJboss,WildFly, IBM Websphere.
Confidential
Java Developer
Responsibilities:
- Involved in the design of the applications using J2EE. This architecture employs a Model/View/Controller (MVC) design pattern.
- Developed code for presentation layer using MVC architecture using Struts framework that uses Servlets and JSP.
- Implemented Action Form class, Action class and Action Mapping for separating the logic from the presentation using Struts.
- Developed presentation layer using HTML, CSS, JSP and JavaScript.
- Developed the helper classes for better data exchange between the MVC layers.
- Coordinating with offshore team, Trouble shootings and solved issues.
- Implemented exception mechanism and used Struts error message mechanism.
- Involved in writing SQL scripts.
- Eclipse is used as an IDE for development.
- Developed ANT Script to compile the Java files and to build the jars and wars.
- Implemented MVC architecture using Struts 1.1 in terms of JSP and Servlets.
- Written JavaScript for validation of page data in the JSP pages.
- Good working knowledge on monitoring and troubleshooting.
Environment: Java,Servlets, JDBC, HTML, CSS, JavaScript, SQL Server, IBM Websphere, JBoss.
