Spark/hadoop Developer Resume
Conshohocken, PA
SUMMARY
- Over 8 years of experience in application development and design using emerging technologies like Hadoop, NoSQL, Big data and Java/J2EE Technologies.
- Around 4+ years of strong working experience with Big Data and Hadoop Ecosystems
- Proven expertise in Machine Learning and Data Science and in using new tools and technologies developments to drive improvements throughout entire software development lifecycle.
- Depth knowledge on Hadoop stack, Data warehousing/Data Mining principles, architecture and development in petabyte scale environments
- Proficiency in Big Data Practices and Technologies like HDFS, MapReduce, Hive, Pig, HBase, Sqoop, Oozie, Flume, Spark, Storm, Impala, Kafka.
- Sophisticated experience with Hadoop distributions like Apache, IBM BigInsights, Cloudera, Hortonworks & MapR.
- Experience in Data Analysis, Data Validation, Data Cleansing, Data Verification and identifying data mismatch.
- Proficient in big data ingestion and streaming tools like Flume, Sqoop, Kafka and Storm.
- Excellent hands on experience in analyzing data using Pig Latin, HQL, HBase and MapReduce programs in Java.
- Extensive hands on experience in writing complex MapReduce jobs, Pig Scripts and Hive data modeling.
- Capable of creating real time data streaming solutions and batch style large scale distributed computing applications using Apache Spark, Spark Streaming, Kafka and Flume.
- Have experience in Apache Spark, Spark Streaming, Spark SQL and No SQL databases like HBase, Cassandra, MongoDB
- Mastery in developing customized UDF’s in Java to extend Hive and Pig Latin functionality.
- Extensively worked on ETL Talend data integration tool.
- Knowledge of administrative tasks such as installing Hadoop (on Ubuntu) and its ecosystem components such as Hive, Pig, Sqoop.
- Accomplished on Big Data Integration and Analytics based on Hadoop, SOLR, Spark, Kafka, Storm and web method technologies
- Utilized Apache Kafka for tracking data ingestion to Hadoop cluster.
- Implemented Kafka Custom encoders for custom input format to load data into Kafka Partitions. Real time streaming the data using Spark with Kafka for faster processing.
- Experience in integrating Hadoop with Apache Storm and Kafka to perform web analytics. Uploaded click stream data from Kafka to HDFS, HBase and Hive by integrating with Storm
- Good knowledge in Linux shell scripting or shell commands.
- Good experience in troubleshooting, performance tuning, optimizing, performance large scale Hadoop cluster with data ingestion, processing and ETL Design/Development/Deployment.
- Insight into importing and exporting data using Sqoop from HDFS to Relational Database Systems and vice - versa.
- Understanding in Apache Flume for efficiently collecting, aggregating, and moving large amounts of log data.
- Good knowledge in designing and implementing end to end Data Security and governance within Hadoop platform using Apache Knox, Apache Sentry, Kerberos etc.
- Proficient in writing the MapReduce jobs on Hadoop ecosystem including Hive and Pig in Cloudera.
- Proficient in job workflow scheduling and monitoring tools like Oozie.
- Strong Database experience with PL/SQL Programming Skills in creating Packages, Stored Procedures, Functions, Triggers & Cursors.
- Practical knowledge with project scoping and planning, risks, issues, schedules and deliverables.
- Exposure to requirements gathering, design and development, application migration and maintenance phases of the Software Development Lifecycle (SDLC).
- Expertise on wide range of middleware technologies like Spring, Spring Integration, Web Services (SOAP and REST services), MQ, EJB (MDB), JAVA, J2EE, XML, XSD etc.
- Expertise in web Technologies like HTML, CSS, PHP, XML, JSP, Ajax, JQuery.
TECHNICAL SKILLS
Big Data/Hadoop: HDFS, Hadoop MapReduce, Hive, Pig, Sqoop, Flume, Oozie, Scala, Storm, Kafka, Talend, Spark, Spark SQL
Languages: C, Java, SQL/PLSQL, Linux Commands
Frameworks: Hadoop, Apache Spark, Apache Kafka, Spring, Restful Web services JAX-RS
Research Interests: Machine Learning, Data Mining, Deep Learning, Real time Processing
Methodologies: Agile, Waterfall model
Web Development: HTML, JavaScript, CSS, Bootstrap, Angular JS, PHP, JEE
Databases: HBase, Cassandra, Oracle, MySQL, SQL server
Web Tools/Frameworks: HTML, CSS, Java Script, XML, JDBC, MVC, Ajax, JSP, Servlets, Struts, Junit, Spring, Hibernate, Tableau
PROFESSIONAL EXPERIENCE
Confidential, Conshohocken PA
Spark/Hadoop Developer
Environment: Hadoop, HDFS, MapReduce, YARN, Spark, Pig, Hive, Sqoop, Flume, Impala, Kafka, HBase, Oozie, Spark, Java, SQL scripting, Linux shell scripting, Eclipse and Cloudera.
Responsibilities:
- Expertise in designing and deployment of Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, Oozie, Sqoop, Flume, Spark, Impala with Cloudera distribution.
- Installed Hadoop, Map Reduce, and HDFS and developed multiple MapReduce jobs in PIG and Hive for data cleaning and pre-processing.
- Assisted in upgrading, configuration and maintenance of various Hadoop infrastructures like Pig, Hive, and HBase.
- Developed workflows and coordinator jobs in Oozie.
- Developed Spark scripts by using Scala shell commands as per the requirement.
- Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
- Developed Scala scripts, UDF’s using both Data frames/SQL and RDD/MapReduce in Spark for Data Aggregation, queries and writing data back into RDBMS through Sqoop.
- Exploring with Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark context, Spark-SQL, Data Frame, pair RDD's, Spark YARN.
- Developed Spark code and Spark-SQL/Streaming for faster testing and processing of data.
- Experience in deploying data from various sources into HDFS and building reports using Tableau.
- Developed a data pipeline using Kafka and Strom to store data into HDFS.
- Performed real time analysis on the incoming data.
- Configured deployed and maintained multi-node Dev and Test Kafka Clusters.
- Performed transformations, cleaning and filtering on imported data using Hive, Map Reduce, and loaded final data into HDFS.
- Load the data into Spark RDD and performed in-memory data computation to generate the output response.
- Loading data into HBase using Bulk Load and Non-bulk load.
Confidential, Philadelphia PA
Hadoop Developer
Environment: Hortonworks Hadoop, HDFS, Hive, Scala, Map Reduce, Storm, Spark, Java, HBase, Pig, Sqoop, Shell Scripts, Oozie Co-ordinator, MySQL, Tableau, Talend, Kafka, SOLR, Storm.
Responsibilities:
- Maintained System integrity of HDFS, HBase, and Hive.
- Maintained the Hadoop cluster and integrated the hive warehouse with HBase
- Involved in migrating the data from MySQL into HDFS using Sqoop and importing various formats of flat files into HDFS.
- Integrated Apache Storm with Kafka to perform web analytics. Uploaded click stream data from Kafka to HDFS, HBase and Hive by integrating with Storm.
- Developed multiple Kafka Producers and Consumers by using low level and high level API's.
- Designed and created Hive external tables using shared meta-store instead of derby with dynamic partitioning and buckets.
- Designed Risk Audit process for a Healthcare client and created Risk assessment database on hive for performing Risk Assessment and Audits and used Tableau for visualization.
- Worked with HiveQL on big data of logs to perform a trend analysis of user behavior on various online modules.
- Supported Map Reduce Programs running on the cluster
- Performed advanced procedures like text analytics and processing, using the in-memory computing capabilities of Spark using Scala.
- Worked on Big Data Integration and Analytics based on Hadoop, SOLR, Spark, Kafka, Storm and web methods technologies.
- Implemented Spark using Scala and Spark SQL for faster testing and processing of data.
- Streaming the data using Spark with Kafka.
- Worked on migrating MapReduce programs into Spark transformations using Scala
- Generate final reporting data using Tableau for testing by connecting to the corresponding Hive tables using Hive ODBC connector.
- Developing and maintaining efficient ETL Talend jobs for Data Ingest.
- Worked on Talend ETL tool, developed and scheduled jobs in Talend integration suite.
- Modified reports and Talend ETL jobs based on the feedback from QA testers and Users in development and staging environments.
- Migrated Hadoop jobs into higher environments like UAT and Production
Confidential, Foster City CA
Hadoop Developer
Environment: Hadoop, Map Reduce, HDFS, Pig, Hive, Sqoop, Flume, Java, Linux, SQL Server.
Responsibilities:
- Extracted and updated the data into HDFS using Sqoop import and export command line utility interface.
- Responsible for developing data pipeline using Flume, Sqoop and pig to extract the data from weblogs and store in HDFS.
- Involved in developing Hive UDFs for the needed functionality.
- Involved in creating Hive tables, loading with data and writing hive queries.
- Utilized Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Applied Pig to do transformations, event joins, filter boot traffic and some pre-aggregations before storing the data onto HDFS.
- Engaged in daily scrum meetings and iterative development.
- Loaded all the data from the existing SQL Server to HDFS using Sqoop.
- Loaded cache data in to HBase using Sqoop.
Confidential, Philadelphia PA
Java /Hadoop Developer
Environment: Java, JSP, HTML, CSS, Java Script, JQuery, Struts 2.0, MySQL, Oracle, Hibernate, JDBC, Eclipse, SQL Stored Procedures, Tomcat, Hive, Pig, Sqoop, Flume.
Responsibilities:
- Design and creation of GUI screens using JSP, Servlets and HTML based on Struts MVC Framework
- Operated JDBC to access Database.
- Manipulated JavaScript for client side validation
- Validations were performed using Struts Validation Framework
- Commit and Rollback methods were provided for transactions processing
- Designed and developed the action form beans and action classes and implemented MVC using Struts framework
- Written Oracle SQL Stored procedures, functions and triggers.
- Developed both Session and Entity beans representing different types of business logic abstractions.
- Maintained the server log document.
- Unit /Integration testing
- Implemented and designed user interface for web based customer application.
- Understanding business needs, analyzing functional specifications and map those to develop and designing MapReduce programs and algorithms.
- Written Pig and Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data. Also have hand on Experience on Pig and Hive User Define Functions (UFD).
- Execution of Hadoop ecosystem and Applications through Apache HUE.
- Optimizing Hadoop MapReduce code, Hive/Pig scripts for better scalability, reliability and performance.
Confidential
Java Programmer
Environment: Eclipse, Java, Ajax, JDBC, JSP, HTML, CSS, XML, JavaScript, MySQL, Tomcat
Responsibilities:
- Designed projects for Client/Server application development.
- Involved in documentation, QA Testing and debugging.
- User Interface development using Eclipse.
- Developed JavaScript API for cross platform application development.
- Utilized Apache Tomcat server integrated with Eclipse for debugging and Unit testing.
- Designed the front end using JSP, HTML, CSS.
- Implemented GUI pages by using JavaScript, HTML, JSP, CSS, and AJAX.
Confidential
Software Engineer
Environment: Eclipse, Java, JSP, Spring, PHP, Ruby, MySQL, HTML, CSS, JavaScript, JQuery
Responsibilities:
- Productized object oriented code, system design and test/QA plans for projects.
- Refined application framework for data flow and data handling.
- Generated the application using Eclipse IDE.
- Implemented Waterfall model practices according to the application requirements.
