Spark Developer Resume
Chicago, IL
SUMMARY
- Has 6 years of IT experience which includes 4 years of work experience in Big Data, Hadoop ecosystem, AWS related technologies.
- Working experience in Apache Hadoop ecosystem components like HDFS, Map Reduce, Pig, Hive, Impala, HBase, SQOOP, Flume, Oozie, Accumulo and Spark.
- Worked and learned a great deal from Amazon Web Services (AWS) Cloud services like EC2, S3, EBS, RDS and VPC.
- Improved the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark - SQL, Data Frame, Pair RDD's, Spark YARN.
- Hands on experience in working on Spark SQL queries, Data frames, import data from Data sources, perform transformations, perform read/write operations, save the results to output directory into HDFS.
- Implemented POC's using Kafka, spark Streaming and Spark SQL.
- Experience in working wif major Hadoop distributions like Cloudera 5.x and Hortonworks HDP 2.x and above.
- Experience in optimizing Map Reduce Programs using combiners, partitioners and custom counters for delivering the best results.
- Experience in writing Pig and Hive scripts and extending the core functionality by writing custom UDF's.
- Good knowledge on File formats such as sequence File, RC, ORC, Parquet and compression techniques such as gzip, snappy and LZO.
- Expert in Tableau data blending and data modeling.
- Designed, developed and implemented Hadoop framework ETL using relational databases and used Tableau reporting on it.
- Integration wif various Hadoop Eco-System Tools:
- Integrated Hive and HBase for better performance
- Integrated Impala and HBase for real-time analytics
- Integrated Hive and Spark SQL for high performance
- Did spark and HBase Integration for OLTP
- Integrated Kafka-Spark streaming for high efficiency throughput and reliability
- Worked on Apache Flume for collecting and aggregating huge amount of log data and stored it on HDFS for doing further analysis.
- Experience in Importing Traditional RDBMS data to HDFS Using Sqoop and Exporting data from HDFS to RDBMS to generate reports.
- Experience in writing both time and data driven workflows using Oozie.
- Solid understanding of algorithms, data structures and object-oriented programming.
- Knowledge on NoSQL columnar databases like HBase and Cassandra.
- Experience in managing and troubleshooting Hadoop related issues.
- Good knowledge and understanding of Java and Scala programming languages.
- Knowledge on Linux and shell scripting.
- Diverse experience in utilizing Java tools in business, Web, client-server platforms using core java, JSP.
- Servlets, Swings, Java Database Connectivity (JDBC) and application servers such as Apache Tomcat.
- Knowledge in using SQL Queries for back-end database analysis.
- Expertise in Agile methodologies.
- Experience in reporting analyzed data in vivid formats using reporting tool Tableau.
- Adept in handling the team in untoward situations and capable of sailing the team to deliver the quality output.
- Highly motivated and versatile team player wif the ability to work independently & adapt quickly to new emerging technologies.
TECHNICAL SKILLS:
Hadoop Ecosystem: MapReduce, Core Java, Apache PIG, Apache HIVE, SQOOP, HBase, Flume, Oozie, MongoDB, Apache Mesos, Scala Impala,Accumulo, YARN, Spark
Reporting Tools: Tableau, Lumira
Domain Knowledge: Retail, Supply Chain, Utilities (IS-U), Pharmaceutical
RDBMS: DB2 UDB, Sybase, Oracle, MS SQL Server and Teradata, MS Access, MySQL
OS: SUSE LINUX, AIX, Sun Solaris UNIX, Vista, Windows 2008/2003/XP
Other Tools: Visual Studio, Visual InterDev, FrontPage and Visio, AB initio
Programming Languages: C, C++, Java, Python
Application Tools/Packages: ERwin 4.1/4.0, MS Project, SQL Navigator 4.1, Oracle SQL Developer
Methodologies: Agile, Waterfall
PROFESSIONAL EXPERIENCE
Spark Developer
Confidential - Chicago, IL
Responsibilities:
- Responsible for design development of Spark SQL Scripts based on Functional Specifications.
- Exploring wif the Spark improving the Performance and Optimization of the existing algorithms in Hadoop.
- Import the data from different sources like HDFS/HBase into Spark RDD. worked wif Spark Context, Spark-SQL, Pair RDD's, Spark YARN.
- Responsible for Performance Tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and Memory tuning.
- Optimizing of existing algorithms in Hadoop using Spark Context, Spark-SQL and Pair RDD's.
- Responsible in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during Ingestion process itself.
- Exploring wif the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark context, Spark - SQL, Data frame, Pair RDD'S, Spark YARN.
- Experienced in Kafka to ingest data into Spark Engine.
- Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
- Developed Kafka producer and consumers, HBase clients, Spark and Hadoop MapReduce jobs along wif components on HDFS, Hive.
- Worked on analyzing Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase database and SQOOP.
- Installed Hadoop, Map Reduce, HDFS, and Developed multiple map reduce jobs in PIG and Hive for data cleaning and pre-processing.
- Used Accumulo to store and retrieve data.
- Coordinated wif business customers to gather business requirements and also interact wif other technical peers to derive Technical requirements and delivered the BRD and TDD documents.
- Extensively involved in Design phase and delivered Design documents.
- Involved in Testing and coordination wif business in User testing.
- Importing and exporting data into HDFS and Hive using SQOOP.
- Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Involved in creating Hive tables, loading wif data and writing hive queries dat will run internally in map reduce way.
- Experienced in defining job flows.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Experienced in managing and reviewing the Hadoop log files.
- Used Pig as ETL tool to do Transformations, even joins and some pre-aggregations before storing the data onto HDFS.
- Experience wif Teradata Tools and Utilities (FastLoad, MultiLoad, BTEQ, FastExport) Advanced SQL (preferably Teradata)
- Load and Transform large sets of structured data.
- Responsible to manage data coming from different sources.
- Involved in creating Hive Tables, loading data and writing Hive queries.
- Utilized Apache Hadoop environment by Cloudera.
- Created Data model for Hive tables.
- Involved in Unit testing and delivered Unit test plans and results documents.
- Exported data from HDFS environment into RDBMS using Sqoop for report generation and visualization purpose.
Environment: Cloudera, Hadoop, Hive, Map Reduce, Hbase, Pig, SQOOP, Mqsql
Big Data Developer
Confidential - Irving, TX
Responsibilities:
- Primary responsibilities include building scalable distributed data solutions using Hadoop ecosystem.
- Datasets will be loaded from two different sources like Oracle, MySQL to HDFS and Hive respectively on daily basis.
- Installed and configured Hive on the Hadoop cluster.
- Worked on Hbase Java API to populate operational Hbase table wif Key value.
- Developed multiple MapReduce jobs in java for data cleaning and preprocessing.
- Developing and running Map-Reduce jobs on YARN and Hadoop clusters to produce daily and monthly reports as per user's need.
- Scheduling and managing jobs on a Hadoop cluster using Oozie work flow.
- Experience in developing multiple MapReduce programs in java for data extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV and other file formats.
- Hands-on experience wif AWS (Amazon Web Services), using Elastic MapReduce (EMR), creating buckets in S3 and storing data in them.
- Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
- Integrated Apache Storm wif Kafka to perform web analytics. Uploaded click stream data from Kafka to HDFS, Hbase and Hive by integrating wif Storm.
- Designed and developed PIG Latin Scripts to process data in a batch to perform trend analysis.
- Developed HIVE scripts for analyst requirements for analysis.
- Developed java code to generate, compare & merge AVRO schema files.
- Developed complex Map reduce streaming jobs using Java language dat are implemented Using Hive and Pig.
- Optimized Map reduce Jobs to use HDFS efficiently by using various compression mechanisms.
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and Extracted the data from MySQL into HDFS using Sqoop
- Automating and scheduling the Sqoop jobs in a timely manner using Unix Shell Scripts.
- Analyzed the data by performing Hive queries (HiveQL) and running Pig Latin scripts to study customer behavior.
- Developed Data Cleaning techniques / UDFs using Pig scripts / Hive QL, Map/Reduce.
- Worked on NoSQL including MongoDB, Cassandra and HBase.
- Continuously monitored and managed the Hadoop Cluster using Cloudera Manager.
Environment: Hadoop, HDFS, Pig, Pig Latin Eclipse, Hive, Map Reduce, Java,HBase, Sqoop, Storm, LINUX, Cloudera, Big Data, Java, My SQL, NoSQL, MongoDB, Cassandra, JSON, XML, CSV.
Big Data Developer
Confidential
Responsibilities:
- Developing and running Map-Reduce jobs on YARN and Hadoop clusters to produce daily and monthly reports as per user's need.
- Debugging/Troubleshoot issues on UDF's in Hive.
- Scheduling and managing jobs on a Hadoop cluster using Oozie work flow.
- Experienced on loading and transforming of large sets of structured, semi structured and unstructured data.
- Transforming unstructured data into structured data using PIG.
- Designed and developed PIG Latin Scripts to process data in a batch to perform trend analysis.
- Good experience on Hadoop tools like MapReduce, Hive and HBase.
- Worked on both External and Managed HIVE tables for optimized performance.
- Developed HIVE scripts for analyst requirements for analysis.
- Worked on Developing custom MapReduce programs and User Defined Functions (UDFs) in Hive to transform the large volumes of data wif respect to business requirement.
- Maintenance of data importing scripts using Hive and Map reduce jobs.
- Data design and analysis in order to handle huge amount of data.
- Cross examining data loaded in Hive table wif the source data in oracle.
- Working close together wif QA and Operations teams to understand, design, develop, end to-end data flow requirements.
- Utilizing oozie to schedule workflows.
- Developing structured, efficient and error free codes for Big Data requirements using my knowledge in Hadoop and its Eco-system.
- Storing, processing and analyzing huge data-set for getting valuable insights from them.
Environment: HDFS, Pig, Hive, HBase, Map Reduce, Sqoop, Oozie, LINUX, Java, SQL
Big data/ Hadoop Admin
Confidential
Responsibilities:
- Installed Cloudera CDH 4 cluster (name node and data nodes).
- Monitored Hadoop cluster and file system usage on name and data nodes.
- Created user accounts and given users the access to the Hadoop cluster.
- Performed HDFS cluster support and maintenance tasks like adding and removing nodes.
- Lead various data conversion initiatives dat included testing Hadoop SQL performance using bucketing and partitioning.
- Installed and configured an automated tool Puppet dat included the installation and configuration of the Puppet master, agent nodes and an admin control workstation.
- Wif Puppet, automated the installation and configuration of the Hadoop Cluster, Hive, HBase, Flume, Pig and Sqoop.
- Use of Sqoop to import and export data from HDFS to RDBMS and vice-versa.
- Implemented HIVE security using Kerberos.
- Exported all log files generated from various sources to HDFS for further processing.
- Monitoring Hadoop Cluster through Cloudera Manager and Implementing alerts based on Error messages.
- Providing reports to management on Cluster Usage Metrics.
- Working wif data delivery teams to setup new Hadoop users. dis job includes setting up Linux users, setting up Kerberos TEMPprincipals and testing HDFS, Hive.
- Worked wif application teams to install operating system, Hadoop updates, patches, version upgrades as required.
- Performed installation of Apache web server for Ganglia to publish charts and graphs in the web console.
- Set up automated monitoring for Hadoop cluster using Ganglia, which halped figure out the load distribution, memory usage and provided an indication for more space.
Environment: Hadoop 2.5, HDFS, Map Reduce, Hive, Flume, Sqoop, Cloudera Manager, CDH4, HBase, Puppet, Pig, AWS EC2.
SQL /BI Developer
Confidential
Responsibilities:
- Responsible for ensuring a stable production environment using change management systems.
- Extensively Wrote and Edited Store Procedure. User Defined Functions and Triggers in T-SQL.
- Participated in the data warehousing of major deployment and Migration projects.
- Monitor database system details wifin the database, including stored procedures and execution time, and implement efficiency improvements.
- Developed tune queries, assigned proper indexes, and aided by developers in the creation of triggers and stored procedures.
- Monitored log space of all the instances wif generation of reports.
- Involved in complete Software Development Life Cycle (SDLC) process by analyzing business requirements and understanding the functional work flow of information from source systems to destination systems.
- Designed and developed SSIS Packages to import and export data from MS Excel, SQL Server 2005 and Flat files.
- Involved in complete SQL 2008 to SQL 2016 SSIS migration and SSIS life cycle in creating SSIS packages, building, deploying and executing the packages in both the environments (Development and Production).
- Used various Transformations in SSIS Dataflow, Control Flow using for loop Containers and Fuzzy Lookups etc.
- Involved in ETL architecture enhancements to increase the performance using query optimizer.
- Created T-SQL queries for data retrieval and created other database objects like tables, views, Stored Procedures, Functions, Triggers and indexes.
- Installed and configured SQL Server successfully.
- Experience in Creating and Updating Clustered and Non-Clustered Indexes to keep up the SQL Server Performance.
- Maintained the table performance by following the tuning tips like normalization, creating indexes and collect statistics.
- Import & Export of data from one server to other servers using SQL Server Integration Services (SSIS).
- Build efficient SSIS packages for processing fact and dimension tables wif complex transforms.
- Created SSIS package to get data from different sources, consolidate and merge into one single source.
- Generated periodic reports based on the statistical analysis of the data from various time frame and division using SQL Server Reporting Services (SSRS).
- Developed various operational, drill-through and drill-down reports using SSRS.
Environment: SQL Server 2012/2008 R2, SSRS, SSIS, XML, MS Excel, MS Access.
