Spark & Hadoop developer Resume
Columbus, OH
SUMMARY:
- IT Professional with 8 years of experience in Software applications development including Analysis, Design, Development, Integration, Testing and Maintenance of various applications using JAVA/J2EE technologies, 4 years of work experience in Bigdata/ Hadoop Development and Ecosystem Analytics using programming languages like Java and Scala.
- Expertise in Big data architecture with Hadoop File system and its eco system tools MapReduce, HBase, Hive, Agile, Pig, Zookeeper, Oozie, Flume, Avro, Impala, Apache spark and Spark Streaming and Spark SQL.
- Experienced in building highly scalable Big - data solutions using Hadoop and multiple distributions i.e., Cloudera, Hortonworks and NoSQL platforms(Hbase& Cassandra).
- Experience in analyzing data using HiveQL, Pig Latin and writing custom mapreduce programs in Java and Python.
- Expert in implementing advanced procedures like text analytics and processing using the in-memory computing capabilities like Apache Spark written in Scala.
- Experienced in converting HiveQL queries into Spark transformations using Spark RDDs and Scala.
- Hands on experience in Apache Sqoop, Apache Storm and Apache Hive integration.
- Hands on experience working with different File Formats like TEXTFILE, JSON, AVROFILE, ORC for HIVE Querying and Processing.
- Experience in Installing, Configuring, Testing Hadoop Ecosystem components, experience on Hadoop clusters using major Hadoop Distributions - Cloudera(CDH3, CDH4 and CDH5) and HortonWorks.
- Experience in AWS cloud environment and on S3 storage and EC2 instances.
- Expereince on Apache Kafka, used for Messaging broker, Log Aggregation and Stream processing.
- Configured Spark Streaming to receive real time data from the Apache Kafka and store the stream data to HDFS using Scala.
- Expertise in migration data from different databases (i.e. Oracle, DB2, Teradata) to HDFS.
- Involved in creating Hive tables, loading with data and writing hive queries which will run internally in MapReduce pattern, used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Experience in importing and exporting data using Sqoop from Relational Database Systems to HDFS and vice-versa, collecting and aggregating large amount of log data using Apache Flume and storing data in HDFS for further analysis and Job/workflow scheduling and monitoring using Oozie.
- Experience in designing and coding web applications using Core Java & Web Technologies- JSP, Servlets and JDBC, full Understanding of utilizing J2EE technology Stack, including Java related frameworks like Spring, ORM Frameworks(Hibernate).
- Experience on setup Hadoop cluster using MS Azure HDInsight as a part of POC.
- Experience in constructing the ETL jobs that will move data to the operational data Store using Talend ETL Platform.
- Good Knowledge on ETL tools like Informatica and Talend.
- Experience in desigining the UserInterfaces using HTML, CSS, JavaScript and JSP.
- Experience in version control tools like Git.
TECHNICAL SKILLS:
BigData Technologies: HDFS, MapReduce, Hive, Pig, Sqoop, Flume, Oozie, Storm, Zookeeper, Kafka, Impala,HCatalog, Apache Spark, Spark Streaming, Spark SQL, Hbase and Cassandra.
Hadoop Distributions: Cloudera (CDH3/CDH4/CDH5), Horton Works.
Operating Systems: Windows XP/7/8, Linux Distro(Ubuntu), Cent OS.
Programming Languages: Java, Scala and Python
Java Technologies: JDBC, Servlets, JSP, Spring and Hibernate
IDE Tools: Eclipse, NetBeans
Web Technologies: HTML, CSS and JavaScript
Scripting Languages: Unix Shell Scripting
Web Services: SOAP, REST, WSDL
RelationalDatabases: DB Oracle, MySQL, Teradata.
NoSql Databases: Hbase and Cassandra
Application Servers: Tomcat, Web Logic, Web Sphere
Reporting/ETL Tools: Informatica,Talend
MS Office Tools: MS WORD, MS EXCEL, MS POWERPOINT, MS VISIO.
Tools: & Utilities: HP Quality Center, Git, Maven.
PROFESSIONAL EXPERIENCE:
Spark & Hadoop Developer
Confidential, Columbus, OH
Responsibilities:
- Evaluated Business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.
- Worked on analyzing Hadoop cluster and different big data analytical and processing tools including Pig, Hive, Spark, Spark Streaming.
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
- Migrating various Hive UDF’s and queries into Spark SQL for faster requests.
- Handled importing of data from various data sources, performed transformations using Hive , MapReduce, loaded data into HDFS and exported the data from HDFS to MYSQL using Sqoop .
- Configured Spark Streaming to receive real time data from the Apache Kafka and store the stream data to HDFS using Scala .
- Hands on experience in Spark and Spark Streaming creating RDD's , applying operations - Transformation and Actions .
- Used HIVE to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Experience in using Apache Kafka for log aggregations.
- Experience on working with different data types like FLATFILES, ORC, AVRO, JSON.
- Developed Talend jobs for reading log files.
- Involved in implementing Cluster for Cassandra to address HBase limitations.
- Provided design recommendations and thought leadership to sponsors/stakeholders that improved review processes and resolved technical problems and suggested some solution.
Environment: MapReduce, HDFS, Hive, Pig, Spark, Spark-Streaming, Spark SQL, Apache Kafka, Sqoop, Java, Scala, CDH4, CDH5, AWS, Eclipse, Oracle, Git, Shell Scripting and Cassandra.
Hadoop Developer
Confidential, CA
Responsibilities:
- Loading the data from the different data sources like (Teradata, DB2, Oracle and Flatfiles) into HDFS using Sqoop and load into Hive tables, which are partitioned.
- Created different PIG Scripts and converted them as shell command to provide aliases for common operation for project business flow.
- Hands on experience in Apache Sqoop, Apache Storm and Apache Hive integration as part of the project implementation.
- Expereince in using Apache Storm to build real-time data integration systems, to analyze clean, normalize, and resolve large amounts of non-unique data points with low latency and high throughput.
- Experience in working on log files using Apache Storm .
- Implemented various Hive queries for analysis and call them from java client engine to run on different nodes.
- Developed Oozie Workflows for daily incremental loads, which gets data from Teradata and then imported into hive tables.
- Developed scripts to bring the log files from FTP Server and then processing it to load into Hive tables.
- Experience in developing Hive UDFs using Java programming language.
- Experience in creating statistics of logs and extracts useful information from the statistics in real-time using Apache Storm.
- Moved data from HDFS to Hbase using Map Reduce and BulkOutputFormat class.
- Experience in Implementing Rack Topology scripts to the Hadoop Cluster.
- Developed Helper class for abstracting Hbase cluster connection act as core toolkit.
- Participated day-to-day meeting, status meeting, and effective communication with team members.
Environment: MapReduce, HDFS, Hive, Pig, Hbase, Apache Storm, HDP, Sqoop, Java, Eclipse, Oracle, Linux, Shell Scripting, Maven, Git.
Hadoop Java Developer
Confidential, Atlanta, GA
Responsibilities:
- Understand the exact requirement of report from the Business groups and users.
- Imported trading and derivatives data in Hadoop Distributed File System using Eco System components MapReduce, Pig, Hive, Sqoop.
- Was part of activity to setup Hadoop ecosystem at developement & QA Environment.
- Managed and reviewed Hadoop Log files.
- Responsible writing PIG Script and Hive queries for data processing.
- Running Sqoop for importing data from Oracle & Other Database.
- Creation of shell script to collect raw logs from different machines.
- Created Partitions in Hive as static and dynamic.
- Implemented Pig Latin scripts using operators such as LOAD, STORE, DUMP, FILTER, DISTINCT, FOREACH, GENERATE, GROUP, COGROUP, ORDER, LIMIT and UNION.
- Defined some PIG UDFs for some functions such as swap, hedging, Speculation and arbitrage.
- Coded MapReduce program to process unstructured logs file.
- Worked on Import and export data into HDFS and Hive using Sqoop.
- Used parameterize Pig Script and optimized script using illustrate and explain.
- Involved in the process of configuring HA, Kerberos security issues and name node failure restoration activity time to time as a part of zero downtime.
- Implemented FAIR Scheduler as well.
- Developed Power Center mappings to extract data from various databases, Flat files and load into data mart using the Informatica .
- Used Spring framework that handles application logic and makes calls to business, make them as Spring Beans.
- Implemented, configured data sources, session factory and used Hibernate Template to integrate Spring framework with Hibernate .
- Developed JUNIT test cases for application unit testing.
- Used SVN as version control to check in the code, created branches and tagged the code in SVN.
- Used RESTFUL Services to interact with the Client by providing the RESTFUL URL mapping.
- Used Log4j framework to log/track application and debugging.
Environment: Hadoop, MapReduce, HDFS, Hive, Pig, Shell Scripting, Sqoop, Java, Eclipse, Spring, Hibernate, SOAP, REST, SVN, Log4j, Informatica.
