Hadoop Developer Resume
Charlotte, NC
SUMMARY:
- Around 7 years of IT experience in Software Development with 5 years’ work experience as Big Data /Hadoop Developer with good knowledge of Hadoop framework.
- Excellent understanding of Hadoop architecture and various components such as HDFS, YARN, High Availability, and MapReduce programming paradigm.
- Hands on experience in installing, configuring, and using Hadoop ecosystem components like Hadoop 2.x, MapReduce 2.x, HDFS, Oozie, Hive, PIG Kafka, Oozie, Zookeeper, Storm, Spark, Sqoop, Flume, HBase.
- Good Exposure on Map Reduce(JAVA), HiveQL, Pig scripting, Spark SQL(Scala/python).
- Experience in building Data pipelines using Kafka and Spark.
- Experience in managing and reviewing Hadoop log files.
- Hands on experience in Import/Export of data using Hadoop Data Management tool Sqoop.
- Development experience in RDBMS like Oracle, MS SQL Server, Teradata and MYSQL.
- Hands on experience in application development using Java, RDBMS, and Linux shell scripting.
- Experience with distributed systems, large - scale non-relational data stores, RDBMS, NoSQL, map-reduce systems, data modeling, database performance, and multi-terabyte data warehouses.
- Experience working with JAVA J2EE, JDBC, Spring MVC.
- Experience in Software Development Life Cycle (Requirements Analysis, Design, Development, Testing, Deployment and Support).
- Excellent interpersonal and communication skills, creative, research-minded, technically competent and result-oriented with problem solving and leadership skills.
TECHNICAL SKILLS:
Distributed File System\ Distributed Programming: HDFS Hadoop 2.6.0\ MapReduce 2.6.x, Pig 0.12, Spark 1.3\
Languages\NoSQLDatabases\: Java (1.6,1.7), Python, Scala, \ HBase 0.98, MongoDBSQL, C, UNIX Shell Scripting, Pig\
Relational Databases \ Hadoop Distribution: Oracle 10g/9i/, MySQL 5.0, SQL Server, \ Cloudera (CDH4,5 CM), Hortonworks \Teradata V2R5\Data Ingestion / ETL tools\SQL On Hadoop and Spark\Kafka 0.8.x, Flume 1.3.x, \ Hive (0.12,1.1.0), Cloudera Impala 2, \Sqoop 1.4.6\ SparkSQL 1.6.0\
Service, Scheduling, Security\ Tools\: Zookeeper 3.3.6, Oozie 4.0.x, Kerberos\ Eclipse, Git, Maven\
Operation Systems\ WebDev. Technologies \: Linux (CentOS, Ubuntu), UNIX, Windows\ CSS3, HTML5, XML, JavaScript,JAVA EE, Spring MVC, Hibernate
PROFESSIONAL EXPERIENCE:
Confidential - Charlotte, NC
Hadoop Developer
Responsibilities:
- Implemented Hadoop framework to capture user navigation across the application to validate the user interface and provide analytic feedback/result to the UI team.
- Loaded data into the cluster from dynamically generated files using Flume and from relational database management systems using Sqoop.
- Pulling the XML and JSON data using REST API’s.
- Monitored Kafka centered messaging system, including architectures and consumer development. Explored RabbitMQ mechanism for delivering high quality data.
- Developed applications using Spark SQL, Spark Streaming.
- Tested Spark Streaming to optimize streaming process and guarantee data quality.
- Loaded the data from Teradata to HDFS using Teradata Hadoop connectors.
- Used Flume to collect, aggregate and store the web log data onto HDFS.
- Wrote Pig scripts to run ETL jobs on the data in HDFS.
- Used Hive to do analysis on the data and identify different correlations.
- Optimizing the Hive queries using Partitioning and Bucketing techniques, for controlling the data distribution.
- Implemented dash boards that internally use hive queries to perform analytics on structured data, Avro and JSON data.
- Experienced in handling AVRO and JSON data in Hive using Hive SerDe’s.
- Worked on importing and exporting data from Oracle and DB2 into HDFS and HIVE using Sqoop.
- Worked on NoSQL databases including HBase and MongoDB. Configured MySQL Database to store Hive metadata.
- Stored and fast update data in HBase, provided key based access to specific data.
- Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
- Automated all the jobs, for pulling data from FTP server to load data into Hive tables, using Oozie workflows.
- Supported Map Reduce Programs those are running on the cluster.
- Involved in migrating the map reduce jobs into Spark Jobs and Used Spark SQL data frames to load structured and semi structured data into Spark Clusters.
- Utilized Agile Scrum Methodology to help manage and organize a team of 4 developers with regular code review sessions.
Environment: Hadoop 2.4.x, HDFS, Map Reduce 2, Flume 1.3.x, Pig 0.12.0, Hive, Spark 1.6.0, HBase, Sqoop 1.4.6, ZooKeeper 3.x, Cloudera CDH 4&5, Oozie, Flat/XML/JSON Files, MongoDB, Python, NoSQL, MYSQL, UNIX Shell Scripting.
Confidential, Fort Worth, TX
Hadoop Developer
Responsibilities:
- Involved in various phases of Software Development Life Cycle.
- Hadoop installation, configuration of multiple nodes in AWS-EC2 & Cloudera platform.
- Developed multiple Map-Reduce jobs in java for data cleansing and preprocessing.
- Used Flume to collect, aggregate and store the web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
- Analyzed the data by performing Hive queries and running Pig scripts to know user behavior.
- Moved Relational Database data using Sqoop into Hive Dynamic partition tables using staging tables.
- Supported Map Reduce programs running on the cluster.
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
- Worked extensively with Sqoop for importing data from Oracle.
- Experienced with different kind of compression techniques like LZO, GZip, and Snappy.
- Created data pipelines using Oozie workflows and scheduled using Ctrl M.
- Optimized existing Oozie workflows by using parallel execution of actions through fork and join functionality.
Environment: Hadoop 2.4, MapReduce, HDFS, Kerberos, Hive, Java 1.6, SQL, Cloudera, Impala, Pig, Sqoop, Oozie, PL/SQL, MySQL 5.0
Confidential, Seattle, WA
Hadoop Developer
Responsibilities:
- Experienced in running Hadoop streaming jobs to process terabytes ofxml format data.
- Load and transform large sets of structured, semi structured andunstructured data.
- Responsible to manage data coming from different sources.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Experienced in analyzing data with Hive and Pig.
- Experienced in definingjob flows.
- Responsible for operational support of Production system.
- Loading log data directly into HDFS using Flume.
- Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
- Jobs management using Fair scheduler.
- Experienced in managing and reviewingHadoop log files.
- Analyzing data with Hive, Pig and Hadoop Streaming.
Environment: Hadoop, Linux, HDFS, Map Reduce, Pig, Hive, Sqoop, Flume, XML.
Confidential
Hadoop Admin/Developer
Responsibilities:
- Apache Hadoop installation, configuration of multiple nodes in AWS-EC2
- Setup and optimize Standalone-System/Pseudo-Distributed/Distributed Clusters
- Build/Tune/Maintain Hive QL and Pig Scripts for reporting purpose
- Help develop MapReduce programs and define job flows
- Manage and review Hadoop log files
- Support/Troubleshoot MapReduce programs running on the cluster
- Load data from Linux/UNIX file system into HDFS
- Create tables, load data, and write queries in Hive
- Develop scripts to automate routine DBA tasks using Linux/UNIX Shell Scripts/Python (i.e. database refresh, backups, monitoring etc.)
- Tune/Modify SQL for batch and online processes
- Monitor cluster using Ambari and optimize system based on job performance and criteria
- Manage cluster through performance tuning and enhancements
Environment: Hortonworks Hadoop (HDFS), MapReduce, AWS, Hive, Java (JDK 1.6), Flat/XML/JSON Files, PSQL, Linux/UNIX Shell Scripting, Python
Confidential
Java Developer
Responsibilities:
- Involved in various phases of Software Development Life Cycle (SDLC).
- Designed and Developed several multi-tiered J2EE application and products as per an Object Oriented Architecture OR SOA standards.
- Developed user interfaces using JSP framework with AJAX, Java Script, HTML, XHTML, and CSS.
- Involved in the design and developing of various modules using CBD Navigator Framework.
- Understand all project requirements as specified in Use Cases, Requirements.
- Actively participated in design and development of the Home Page, Investment Products and user maintenance screens for internal admin in IAM and FAS application as per UI prototypes.
- Used Hibernate framework to communicate with the DB2 database for various modules.
- Used JavaScript for client side validations and used JSF frame work for server side validation.
- Used XSLT to transform XML documents into HTML templates.
- Performed key role in designing and developing enterprise J2EE applications using RAD.
- Used JUnit Framework for performance testing, code coverage.
- Used J2EE design patterns like Spring MVC
- Actively participated in the Agile Development Process.
Environment: JAVA J2EE, JSP, AJAX, Spring MVC, Hibernate, HTML, CSS, JavaScript, XML, XSLT, Junit, RAD, DB2
