Spark Developer Resume
Houston, TexaS
SUMMARY:
- Around seven years of work experience in IT, which includes experience in Installation, Development and Implementation of Hadoop and Data warehousing solutions.
- Experience in dealing with Apache Hadoop components like HDFS, MapReduce, HiveQL, HBase, Pig, Sqoop, Big Data and Big Data Analytics.
- Hands on experience in MapReduce jobs. Experience in installing, configuring and administrating the Hadoop Cluster of Major Hadoop Distributions.
- Hands on experience in installing, configuring and using echo system components like Hadoop, MapReduce, HDFS, Hbase, Zoo keeper, Hive, Sqoop and Pig.
- Experience developing Kafka producers and Kafka Consumers for streaming millions of events per second on streaming data
- Significant experience writing custom UDF’s in Hive and custom Input Formats in MapReduce.
- Good Knowledge in writing Spark Applications in Scala
- Hands on Experience in designing and developing applications I n Spark using Scala to compare the performance of Spark with Hive and SQL/Oracle.
- Good Knowledge in Spark - Data Frames, Data sets.
- Expertise in writing the Real - time processing application Using spout and bolt in Storm .
- Knowledge of job workflow management and coordinating tools like Oozie.
- Strong experience productionailing end to end data pipelines on Hadoop platform.
- Software developer in Java Application Development, Client/Server Applications, and Internet/Intranet based database applications and developing, testing and implementing application environment using J2EE, JDBC, JSP, Servlets, Web Services, Oracle, PL/SQL and Relational Databases.
- Experience in design, development of web-based applications using HTML, DHTML, CSS, JavaScript, JQuery, JSP and Servlets.
- Good work experience with Hibernate open source object/relational mapping framework.
- Solid design skills using Java Design Patterns and Unified Modeling Language UML.
- Experience in working with different operating systems Windows 98/NT/2000/XP/2007/2008, UNIX, and LINUX.
- Experience working with NoSQL database technologies, including MongoDB, Cassandra and HBase.
- Extensive experience in importing/ exporting data to / from RDBMS and HDFS using Apache Sqoop.
- Expertise in working with HIVE data warehouse infrastructure-creating tables, data distribution by implementing Partitioning and Bucketing, developing and tuning the HQL queries.
- Involved in creating Hive tables, loading with data and writing Hive queries that will run internally in MapReduce and TEZ.
- Strong understanding of real time streaming technologies Spark and Kafka.
- Strong understanding of Logical and Physical data base models and entity-relationship modeling.
- Replaced existing MR jobs and Hive scripts with Spark SQL & Spark data transformations for efficient data processing.
- Good understanding of the Data modeling (Dimensional & Relational) concepts like Star-Schema Modeling,
- Snowflake Schema Modeling, Fact and Dimension tables.
- Experience in manipulating/analyzing large datasets and finding patterns and insights within structured and unstructured data.
- Strong understanding of Java Virtual Machines and multithreading process.
- Experience in writing complex SQL queries, creating reports and dashboards.
- Expertise in handling ETL tools like Informatica, Talend.
- Excellent analytical, communication and interpersonal skills.
- Possess excellent communication, interpersonal and analytical skills along with positive attitude.
TECHNICAL SKILLS:
Programming/Scripting Languages: Scala, Java, Python, C, C++, SQL
Big Data: Hadoop, MapReduce, HDFS, Hive, Pig and Sqoop,Spark 2.0
Other tools: Microsoft Office tools, VSTS, VM ware, IoT, Git
NoSQL,OracleSQL,MSQL,RDBMS,Apache: Cassandra, Hbase
Bigdata Eco System: HDFS, Oozie, Zookeeper, Spark storm, Spark streaming
Machine Learning: Decision Tree, Neural Networks, ANN & RNN, PCA, Clusters, SVM, K-NN, Deep learning
UNIX Tools: Apache, Yum, RPM
FILE FORMATS: Txt, XML, JSON, Avro, Parquet, ORC
Cloud Computing: AWS
Visualization and Reporting Tools: Tableau, Microsoft Power BI
Methodologies: Agile, UML, Design Patterns
PROFESSIONAL EXPERIENCE:
Confidential - Houston, Texas
Spark Developer
Responsibilities:
- Worked with Hadoop 2.x version and Spark 2.x (Python and Scala).
- Used Spark for interactive queries, processing of streaming data and integration with NoSQL database for huge volume of data.
- Experienced in handling large datasets using Partitions, Spark in-memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself.
- Developed custom ETL solutions, batch processing and real-time data ingestion pipeline to move data in and out of Hadoop using Python and shell scripting.
- Explored with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Pair RDD's, Spark YARN.
- Worked with Sqoop import and export functionalities to handle large data set transfer between Oracle databases and HDFS.
- Developed Spark jobs to clean data obtained from various feeds to make it suitable for ingestion into Hive tables for analysis.
- Imported data from various sources into Spark RDD for analysis.
- Developed Custom Input Formats in MapReduce jobs to handle custom file formats.
- Configured Oozie workflow to run multiple Hive and Pig jobs which run independently with time and data availability.
- Utilized Hive tables and HQL queries for daily and weekly reports. Worked on complex data types in Hive like Structs and Maps.
- Developed PigLatin scripts and HiveQL queries for trend analysis and pattern recognition on user data.
- Extensively used Pig for cleansing and pre-process the data for analysis.
- Helped this regional bank streamline business processes by developing, installing and configuring Hadoop ecosystem components that moved data from individual servers to HDFS.
- Imported data from AWS S3 into Spark RDD, performed transformations and actions on RDDs.
- Worked with Hadoop distribution of Horton Works
- Installed and configured MapReduce, HIVE and the HDFS; implemented CDH3 Hadoop cluster on CentOS. Assisted with performance tuning and monitoring.
- Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
- Supported code/design analysis, strategy development and project planning.
- Created reports for the BI team using Sqoop to export data into HDFS and Hive.
- Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
- Assisted with data capacity planning and node forecasting.
- Collaborated with the infrastructure, network, database, application and BI teams to ensure data quality and availability.
- Designing ETL processes using Informatica to load data from Flat Files, Oracle and Excel files to target Oracle Data Warehouse database.
- Administrator for Pig, Hive and Hbase installing updates, patches and upgrades.
Environment: Spark,Spark SQL,MapReduce, HBase, Hive, Oozie,Informatica, HQL, Sqoop, Flume, Oozie, Java,Scala, Python, Shell scripting, Maven, Eclipse, Putty, GIT, Tableau
Confidential - Columbus, Ohio
Hadoop developer
Responsibilities:
- Involved in review of functional and non-functional requirements.
- Installed and configured Pig and also written PigLatin scripts.
- Wrote MapReduce job using Pig Latin. Designed and implemented an ETL framework using Java and PIG to load data from multiple sources into Hive and from Hive into Vertica.
- Utilized SQOOP, Kafka, Flume and Hadoop File System API’s for implementing data ingestion pipelines.
- Worked on real time streaming, performed transformations on the data using Kafka and Spark Streaming.
- Migrated existing on-premise application to AWS and used AWS services like EC2 and S3 to process and store small data sets.
- Experienced in maintaining the Hadoop cluster on AWS EMR.
- Worked on importing metadata into Hive and migrating existing tables and applications to work on Hive and AWS cloud.
- Involved in managing and reviewing Hadoop log files. Designed and implemented an ETL framework using Java and PIG to load data from multiple sources into Hive and from Hive into Vertica.
- Utilized SQOOP, Kafka, Flume and Hadoop File System API’s for implementing data ingestion pipelines.
- Worked on real time streaming, performed transformations on the data using Kafka and Spark Streaming.
- Developed Spark Streaming jobs in Scala to consume data from Kafka Topics, made transformations on data and insert into HBase tables.
- Worked on converting Hive/SQL queries into Spark transformations using Spark RDDs, Python, and OOP with Python. Worked on developing and executing shell scripts to automate the jobs.
- Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
- Developing Scripts and Batch Job to schedule various Hadoop Program.
- Written Hive queries for data analysis to meet the business requirements.
- Creating Hive tables and working on them using Hive QL.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Experienced in defining job flows.
- Got good experience with NOSQL database SOLR HBase.
- Involved in creating Hive tables loading with data and writing hive queries which will run internally in map reduce way.
- Developed a custom Filesystem plug in for Hadoop so it can access files on Data Platform.
- This plugin allows Hadoop MapReduce programs HBase Pig and Hive to work unmodified and access files directly.
- Designed and implemented MapReduce-based large-scale parallel relation-learning system
- Extracted feeds form social media sites such as Facebook Twitter using Python scripts.
- Setup and benchmarked Hadoop/HBase clusters for internal use
Environment: Hadoop, MapReduce, HDFS, Hive, Java, Hadoop distribution of Horton Works, Cloudera, Pig, HBase, Linux, XML, MySQL,Hadoop, HDFS, ETL, Vertica, Kafka, YARN, Drill, Pig, Hive, NiFi.
Confidential, NJ
Hadoop Developer
Responsibilities:
- Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the EDW.
- Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with EDW reference tables and historical metrics.
- Enabled speedy reviews and first mover advantages by using Oozie to automate data loading into the Hadoop Distributed File System and PIG to pre-process the data.
- Performed ETL operations using Pig, Hive to transform transactional data into de-normalized form.
- Created adhoc reports by gathering requirements from different teams.
- Utilized Pig and Hive user defined functions to analyze the complex data to find specific user behavior.
- Analyzed data using HiveQL to derive metrics like game duration, daily active users(DAU), weekly active users (WAU) etc.
- Implemented Hive generic UDFs to in corporate business logic into Hive queries.
- Worked on joining raw data with the reference data using Pig scripts.
- Worked along with the admin team to assist them in adding/ removing cluster nodes, cluster monitoring and trouble shooting.
- Exported data to relational databases using Sqoop for visualization and to generate reports.
- Created Machine Learning and statistical models like (SVM, CRF, HMM) to assess gamer performance.
- Provided design recommendations and thought leadership to sponsors/stakeholders that improved review processes and resolved technical problems.
- Managed and reviewed Hadoop log files.
- Tested raw data and executed performance scripts.
- Shared responsibility for administration of Hadoop, Hive and Pig.
Environment: Hadoop, MapReduce, HDFS, Hive, Pig, Java, SQL, PL/ SQL, MySQL.Sqoop, Linux XML MySQL.
Confidential
Java/J2EE Developer
Responsibilities:
- Install, configure and deploy software, provide quality assurance.
- Troubleshoot various software issues using debugging process and coding techniques.
- Provide high-level customer support to remote clients using a support e-ticketing system.
- Perform system administration for hosting server and client software.
- Developed JSPs and Servlet.
- Developed screens using Java, HTML, DHTML, CSS, JSP and JavaScript.
- Designed Database for the application.
- Implemented all validations and done testing.
- Implemented and managed SQL database for use in background for security and internal proprietary processes.
- Diagnose and correct errors within Java/HTML/PHP code to allow for connection and utilization of proprietary applications.
- End user support and administrative functions to include password and account management.
- Developed PL/SQL View function in Oracle 9i database for get available date module.
- Used Quartz schedulers to run the jobs in a sequential with in the given time
- Used JSP and JSTL Tag Libraries for developing User Interface components.
Environment: JDK1.2, JavaScript, HTML, DHTML, XML, Struts, JSP, Servlet, JNDI, J2EE, Tomcat, Oracle, JSP.
Confidential
Java/J2EE Developer
Responsibilities:
- Analyzed and reviewed client requirements and design
- Worked on testing, debugging and troubleshooting all types of technical issues.
- Good knowledge in OOPS concepts, OOAD, UML
- Used JDBC for database connectivity and manipulation
- Used Eclipse for the Development, Testing and Debugging of the application.
- Working as java j2ee backend developer in creating the Maven web application project
- Used DOM Parser to parse the xml files.
- Log4j framework has been used for logging debug, info & error data.
- Used WinSCP to transfer file from local system to other system.
- Performed Test Driven Development (TDD) using JUnit.
- Used JProfiler for performance tuning
- Built the application using MAVEN and deployed using WebSphere Application server
- Gathered and collected information from various programs, analyzed time requirements and prepared documentation to change existing programs.
- Used SOAP for exchanging XML based messages.
- Used Microsoft VISIO for developing Use Case Diagrams, Sequence Diagrams and Class Diagrams in the design phase.
- Developed Custom Tags to simplify the JSP code.
- Designed UI screens using JSP and HTML.
Environment: Java, HTML,Servlets, JSP, Hibernate, Junit Testing, Oracle DB, SQL, Jasper Reports, iReport, Maven, Jenkins.
