Sr. Hadoop/spark developer Resume
Woburn, MA
SUMMARY:
- Hadoop / Spark Developer with8+ years of involvement in Software development which includes 4+ years of experience in Big data and Hadoop Ecosystem components and 4 years in Java development.
- Solid subjective knowledge and hands - on experience in dealing with Apache Hadoop components like HDFS, MapReduce, HiveQL, HBase, Pig, Hive, Sqoop, Oozie, Cassandra, Flume, Apache Spark.
- Currently working on Spark and Spark Streamingextensively using Scala as the main programming dialect.
- Experienced working with Spark Streaming , SparkSQL and Kafka for real-time data processing.
- Extensive experience in working with various distributions of Hadoop Enterprise versions of Cloudera (CDH4/CDH5), Hortonworks and good knowledge on Amazon'sEMR (Elastic MapReduce).
- Designing and implementing complete end-to-end Hadoop Infrastructure including Pig , Hive , Sqoop , Oozie , Flume and Zookeeper .
- Hands on expertise in working and designing of Row keys & Schema Design with NOSQL databases like MongoDB 3.0.1, HBase , Cassandra and DynamoDB (AWS) .
- Designing and creating Hive external tables using shared meta-store instead of derby with partitioning, dynamic partitioning and buckets.
- Extensively used Spark Data frames, Spark-SQL and RDD API of Spark for performing various data transformations and dataset building.
- Experience in importing and exporting data using Sqoop from Relational Databases to HDFS and vice-versa.
- Exposure to Data Lake Implementation using Apache Spark and developed Data pipe lines.
- Extensively worked on Spark using Scala on cluster for computational (analytics), installed it on top of Hadoop performed advanced analytical application by making use of Spark with Hive and SQL/Oracle.
- Extensive experience in importing and exporting streaming data into HDFS using stream processing platforms like Flume and Kafka messaging system.
- Strong experience and knowledge of real time data analytics using Spark Streaming , Kafka and Flume .
- Experience in developing data pipeline using Sqoop , and Flume to extract the data from weblogs and store in HDFS . Accomplished developing Pig Latin Scripts and using Hive Query Language for data analytics.
- Working knowledge of Amazon'sElasticCloudCompute (EC2) infrastructure for computational tasks and Simple Storage Service ( S3 ) as Storage mechanism.
- Experienced in migrating data from various sources using PUB-SUB model in Apache Kafka , and Kafka producers, consumers and preprocess data using Storm topologies.
- Developed customized UDFs and UDAFs in java to extend Pig and Hive core functionality.
- Excellent understanding and knowledge of NOSQL databases like HBase and Cassandra.
- Experienced in developing Cassandra data model and administering the Cassandra Hadoop Cluster along with pig and Hive
- Expertise in designing columnar families in Cassandra and writing queries in CQL to analyze data from Cassandra tables.
- Experience in using DataStax Spark-Cassandra connectors to get data from Cassandra tables and process them using Apache Spark.
- Expert knowledge in setting up MongoDB clusters and handling service requests for MongoDB.
- Monitoring of Document growth and estimating storage size for large MongoDB clusters.
- Experienced in creating HBase tables and column families to store the user event data and wrote automated HBase test cases for data quality checks using HBase command line tools.
- Experienced in developing end to end data processing pipelines that begin with receiving data using distributed messaging systems Kafka through persistence of data into HBase .
- Experienced working with Apache Nifi for building and automating dataflow from data source to HDFS and HDFS to Teradata.
- Extensive experience in ETL Architecture, Development, enhancement, maintenance, Productionsupport, Data Modeling, Data profiling, Reporting including Business requirement, systemrequirement gathering.
- Hands-on experience in Shellscripting. Knowledge on cloud services Amazonwebservices(AWS).
- Proficient in using RDMS concepts with Oracle, SQLServer and MySQL.
- Experience in processing different file formats like XML, JSON and sequence file formats.
- Good Knowledge in AmazonAWS concepts like EMR and EC2 web services which provides fast and efficient processing of Big Data.
- Good Experience in creating Business Intelligence solutions and designing ETL workflows using Tableau.
- Knowledge on Enterprise Data Warehouse (EDW) architecture and various data modeling concepts like star schema, Snowflake schema and Teradata.
- Good working experience on different OS like UNIX/Linux, Apple Mac OS-X Windows.
- Experience working both independently and collaboratively to solve problems and deliver high quality results in a fast-paced, unstructured environment.
TECHNICAL SKILLS:
Hadoop Ecosystems: HDFS, Map Reduce, Pig, Hive, Sqoop, Flume, YARN, Oozie, Zookeeper, Impala, Ambari, Spark, Spark SQL, Spark Streaming, Apache Kafka, Storm.
Languages: C, C++, Java, Scala, Python, C#, SQL, PL/SQL, Pig Latin, HiveQL
Frameworks: J2EE, Spring, Hibernate, Angular JS
Web Technologies: HTML, CSS, Java script, jQuery, Ajax, XML, SOAP, REST API, ASP .Net
NoSQL: HBase, Cassandra, MongoDB, DynamoDB
Cluster Management and Monitoring: Cloudera Manager, Hortonworks Ambari, Apache Mesos
Relational Databases: Oracle 11g, MySQL, SQL-Server, Teradata
Build Tools: ANT, Maven, SBT, Jenkins
Application Server: Tomcat 6.0, WebSphere 7.0
Version Control: GitHub, Bit Bucket, SVN
Security: Kerberos, OAuth
Development Methodologies: Agile, Scrum, Waterfall.
PROFESSIONAL EXPERIENCE:
Sr. Hadoop/Spark Developer
Confidential - Woburn, MA
Roles & Responsibilities:
- Worked with Hortonworks distribution of Hadoop for setting up the cluster and monitored it using Ambari.
- Created ODBC connection through Sqoop between Hortonworks and SQL Server.
- Worked with ELK Stack cluster for importing logs into Logstash, sending them to Elasticsearch nodes and creating visualizations in Kibana.
- Used ApacheSpark with ELK cluster for obtaining some specific visualizations which require more complex data processing/querying.
- Developed Spark Applications by using Scala, Java and Implemented Apache Spark data processing project to handle data from various RDBMS and Streaming sources.
- Worked with the Spark for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Spark MLlib, Data Frame, Pair RDD's, Spark YARN.
- Experience in implementing Spark RDD's in Scala.
- ImplementedSpark SQL for faster processing of data and handle Skew Data for real time analysis in Spark.
- Implemented Spark using Scala and utilizing Dataframes and Spark SQL API for faster processing of data.
- Written transformations in Apache Spark using Data frames, Scala and Spark SQL.
- Configured Spark streaming to get streaming information from the Kafka and store them in HDFS.
- Migrated Flume with Spark for real time data and Developed the SparkStreaming Application with java to consume the data from Kafka and push them into Hive.
- Used Kafkafunctionalities like distribution, partition, replicated commit log service for messaging systems by maintaining feeds.
- Developed a NiFi Workflow to pick up the data fromData Lake and from SFTP server and send that to Kafka broker.
- Involved in loading data from rest endpoints to KafkaProducers and transferring the data to KafkaBrokers.
- Used ApacheKafka to aggregate web log data from multiple servers and make them available in Downstream systems for Data analysis and engineering type of roles.
- Used Apache Zookeeper for configuration management and cluster coordination services.
- Developed Preprocessing job using Spark Data frames to flatten JSON documents to flat file.
- Load D-Stream data into Spark RDD and do in memory data Computation to generate Output response.
- Involved in performance tuning of Spark jobs using Cache and complete advantage of cluster environment.
- Designed Columnar families in Cassandra and Ingested data from RDBMS, performed data transformations, and then exported the transformed data to Cassandra as per the business requirement.
- Tested the cluster Performance using Cassandra-stress tool to measure and improve the Read/Writes.
- Developed Sqoop Jobs to load data from RDBMS, External Systems into HDFS and HIVE.
- Developed Oozie coordinators to schedule Pig and Hive scripts to create Datapipelines.
- Written several Map reduce Jobs using Java API, also Used Jenkins for Continuous integration
- Imported data from AWSS3 into Spark RDD, Performed transformations and actions on RDD's.
- Implemented ElasticSearch on Hive data warehouse platform.
- Used Hive QL to analyze the partitioned and bucketed data, executed Hive queries on Parquet tables stored in Hive to perform data analysis to meet the business specification logic.
- Implemented ETL standards utilizing proven data processing patterns, migrated tools from Informatica to Talend
Environment: Hadoop, Spark, Spark-Streaming, Spark SQL, AWS, HDFS, Hive, Pig, Apache Kafka, Sqoop, Java (JDK SE 6, 7), Scala, Shell scripting, Linux, MySQL, Jenkins, Oracle, Oozie, MySQL, NIFI, Cassandra.
Sr. Hadoop Developer
Confidential - Norwalk, CT
Roles & Responsibilities:
- Involved in Installation, Configuring and managing Hadoop cluster using the ClouderadistributionCDH 5.0 and Continuous monitoring of the Hadoop cluster using the Clouderamanager.
- Using Kafka and Kafkabrokers we initiatedspark context and processed live streaming information with the help of RDD as is.
- Storing schema of incoming data sources in Schema registry of Kafka which will be utilized by the downstream applications.
- Real time processing of raw data stored in Kafka and storing processed data in Hadoop using Spark Streaming(DStreams).
- Responsible for loading Data pipelines from webserversusing Sqoop with Kafka and SparkStreamingAPI.
- In pre-processing phase used SparkRDD transformations to remove all the missing data and to create new features.
- Developed SparkSQL queries for generating statistical summary and filtering/aggregation operations for specific use cases working with SparkRDD's on distributed cluster running ApacheSpark.
- Involved in converting SQL queries into Apache Spark transformations using ApacheSparkDataFrames.
- As a part of Data acquisition, used Sqoop and flume to inject the data from server to Hadoop using incremental import.
- Configured SparkStreaming to receive real time data from the Kafka and store the stream data to HDFS.
- Installed, configured and managed Cassandra database and performed read / writes using Java JDBC connectivity.
- Ingested data from Relational databases like MYSQL and Oracle DB2 into HDFS using Sqoop and ingesting them into Cassandraand performed data transformations, and then export the transformed data to Cassandraas per the business requirement.
- Used Sqoop to import the data from databases to Hadoop Distributed File System (HDFS) and performed automated data auditing to validate the accuracy of the loads.
- Developed Oozie Scheduler jobs for performing daily imports using Sqoop incremental imports from Relational databases that store data from upstream servers.
- Configuring, implementing and supporting High Availability (Replication) with Load balancing (sharding) cluster of MongoDB having TB's of data.
- Used Solr Search engine for performing full text searches with MongoDB as a data store.
- Created near Real Time Solr indexing on MongoDB(using Mongo connector from Mongo Labs) and HDFS using Solr Hadoop connector.
Environment: Apache Hadoop, HDFS, Cloudera, Sqoop, Apache Kafka, Oozie, SQL, Scala, Spark, Cassandra, MongoDB, Solr.
Hadoop Developer
Confidential - Gulfport, MS
Roles & Responsibilities:
- Worked on analyzing Hadoop cluster using different big data analytic tools including Pig, Hive, Sqoop, Kafka and impala with Cloudera distribution.
- Involved in complete Implementation lifecycle, specialized in writing customPig and Hive queries.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Experienced in using Impala for faster processing of large datasets from Hadoop clusters and integrated it with BI tools to run ad-hoc queries directly on Hadoop .
- Involved in developing Impala scripts for extraction, transformation, loading of data into data warehouse.
- Experience in developing customized UDF's in java to extend Hive and Pig Latin functionality.
- Created HBase tables to store various data formats of data coming from different sources.
- Managing and scheduling Jobs to remove the duplicate log data files in HDFS using Oozie.
- Developed Flume ETL job for ingesting log data from HTTP Source to HDFS sink.
- Used Flume extensively in gathering and moving log data files from Application Servers to HDFS.
- Implemented test scripts to support test driven development and continuous integration.
- Dumped the data from HDFS to MYSQL database and vice-versa using Sqoop.
- Used File System check (FSCK) to check the health of files in HDFS.
- Developed the UNIX shell scripts for creating the reports from Hive data.
- Involved in the pilot of Hadoop cluster hosted on Amazon Web Services (AWS).
- Extensively used Sqoop to get data from RDBMS sources like Teradata and Netezza.
- Extracted files from MongoDB through Sqoop and placed in HDFS for processing.
- Ingested data from RDBMS and performed data transformations, and then export the transformed data to MongoDB as per the business requirement.
- Experience with creating script for data modeling and data import and export. Extensive experience in deploying, managing and developing MongoDBclusters.
Environment: Apache Hadoop, Map Reduce, HDFS, Ambari, Hive, Sqoop, Apache Kafka, Oozie, SQL, Flume, Scala, Spark, Java, AWS, GitHub.
Hadoop Developer
Confidential - Mobile, AL
Roles & Responsibilities:
- Installed and configured Hadoop Ecosystem components and Cloudera manager using CDH distribution .
- Developed multiple Map Reduce jobs in Java for data cleansing and preprocessing.
- Developed Sqoop scripts to import/export data from Oracle to HDFS and into Hive tables.
- Worked on collecting and aggregating substantial amounts of log data using Flume and staging data in HDFS.
- Worked on analyzing Hadoop clusters using Big Data Analytic tools including Map Reduce , Pig and Hive .
- Involved in creating tables in Hive and writing Hive queries (HQL) to load data into Hive tables from HDFS.
- Optimized the Hive tables using partitions and bucketing to give better performance for Hive QL queries.
- Worked on Hive/Hbase vs RDBMS, imported data to hive, created internal and external tables, partitions, indexes, views, queries and reports for BI data analysis.
- Developed Java custom record reader, partition and serialization techniques.
- Developed interactive shell scripts for scheduling various data cleansing and data loading process
- Used different data formats (Text format and Avro format) while loading the data into HDFS.
- Created tables in HBase and loading data into HBase tables.
- Developed scripts to load data from HBase to Hive Meta store and perform Map Reduce jobs.
- Developed Custom Loaders and Storage Classes in PIG to work on several data formats like JSON, XML, CSV and generated Bags for processing using pig.
- Created custom UDF's in Pig and Hive .
- Created partitioned tables and loaded data using both static partition and dynamic partition methods.
- Installed Oozie workflow engine and scheduled it to run data/time dependent Hive and Pig jobs
- Designed and developed Dashboards for Analytical purposes using Tableau .
- Analyzed the Hadoop log files using Pig scripts to oversee the errors.
Environment: HDFS, Map Reduce, Hive, Sqoop, Pig, HBase, Oozie, CDH distribution, Java, Eclipse, Shell Scripts, Tableau, Windows, Linux.
Java Developer
Confidential
Roles & Responsibilities:
- Understanding requirement and the technical aspects and architecture of the existing system
- Help Design application development using Spring MVC framework, front-end interactive page design using HTML, JSP, JSTL, CSS, JavaScript, jQuery and AJAX.
- Utilized various JavaScript and jQuery libraries, AJAX for form validation and other interactive features.
- Involved in writing SQL queries using SQL query builder for fetching data from Oracle database.
- Designed and developed Web Services to store and retrieve user profile information from database.
- Used Spring DAO concept to interact with Database using JDBC template and Hibernate template.
- Well Experienced in deploying and configuring applications onto application servers like Web logic, WebSphere and Apache Tomcat.
- Created RESTful web services interface to Java-based runtime engine and accounts.
- Performed code walk through for the team members to check the functional coverage and coding standards.
- Followed AGILE Methodology and SCRUM to deliver the product with cross-functional skills.
- Used JUnit to test persistence and service tiers. Involved in unit test case preparation.
- Hands on experience in software configuration and version control tools Subversion (SVN),Gitand CVS.
- Actively used the defect tracking tool JIRA to create and track the defects during QA phase of the project
- Involved in sprint planning, code review and daily standup meetings to discuss the progress of the application.
Environment: Java/J2EE, HTML, Ajax, Servlets, JSP, SQL, JavaScript, CSS, XML, Windows, Unix, Tomcat Server, Spring MVC, Hibernate, JDBC, Git, SVN.
Java Developer
Confidential
Roles & Responsibilities:
- Involving in Analysis, Design, Implementation and Bug Fixing Activities.
- Involving in Functional & Technical Specification documents review.
- Created and configured domains in production, development and testing environments using configuration wizard.
- Deployed and tested the application using Tomcat web server.
- Analysis of the specifications provided by the clients and involved in Application designing.
- Developed Use Case Diagrams, Class Diagrams, Sequence Diagram, Data Flow Diagram.
- Web related development with JSP, AJAX, HTML, XML, XSLT, and CSS.
- Create and enhance the stored procedures, PL/SQL, SQL for Oracle 9i RDBMS.
- Designed and implemented a generic parser framework using SAX parser to parse XML documents.
- Deployed the application on WebLogic Application Server 9.0.
- Extensively used UNIX /FTP for shell Scripting and pulling the Logs from the Server.
- Provided further Maintenance and support, this involves working with the Client and solving their problems which include major Bug fixing.
Environment: Java 1.4, Web logic Server 9.0, Oracle 10g, Web services Monitoring, Web Drive, Unix /Linux, Web Logic Server, JavaScript, HTML, CSS, XML.
