Sr. Hadoop Developer Resume
Denver, CO
SUMMARY:
- Professional experience of 8+ years in IT which includes 4 years of comprehensive experience in working with Apache Hadoop Ecosystem - components, Spark streaming .
- Over 4+ Years of development experience in Big Data Hadoop Ecosystem components and related tools with data ingestion, importing, exporting, storage, querying, pre-processing and analysing of big data.
- Good Working Expertise on handling Terabytes of structured and unstructured data on huge Cluster environment.
- Experience in using SDLC methodologies like Waterfall, Agile Scrum, and TDD for design and development.
- Expertise in implementing Spark modules and tuning its performance.
- Experienced in performance tuning of Spark applications using various resource allocation techniques and transformations reducing the Shuffles and increasing the Data Locality configurations .
- Expertise in Kerberos Security Implementation and securing the cluster.
- Expertise in creating Hive Internal/External Tables/Views using shared Meta store, writing scripts in Havel and experience in data transformation & file processing, building analytics using Pig Latin Scripts.
- Expertise in writing custom UDFs in Pig & Hive Core Functionality.
- Developed, deployed and supported several Map Reduce applications in Java to handle different types of data.
- Worked with various compression techniques like Avro, Snappy, and LZO.
- Hands on experience dealing with AVRO and Parquet file format, following best Practices and improving the performance using Partitioning, Bucketing, and Map side-joins and creating Indexes .
- Expert in implementing advanced procedures like text analytics, processing and implementing streaming APIs using the in-memory computing capabilities like Apache Spark written in Scala, Python and Scala.
- Experience in Data Load Management, importing and exporting data from HDFS to Relational and non- Relational Database Systems using Sqoop, Flume and Apache Nifi by efficient column mappings and maintaining the uniformity .
- Exported data to various Databases like Teradata (Sales Data Warehouse), SQL-Server, Cassandra using Sqoop.
- Experienced in creating shell scripts to push data loads from various sources from the edge nodes onto the HDFS.
- Experienced in performing code reviews, involved closely in smoke testing sessions, retrospective sessions.
- Experience in scheduling and monitoring jobs using Oozie and Crontab.
- Experienced in Microsoft Business Intelligence tools, developing SSIS (Integration Service), SSAS (Analysis Service) and SSRS (Reporting Service) , building Key Performance Indicators and OLAP cubes.
- Have hands on experience in creating cubes for various reports for end clients, Configured Data Source and Data Source Views, Dimensions, Cubes, Measures, Partitions, KPI’s and MDX Queries .
- Have good exposure with the star, snow flake schema, data modelling and work with different data warehouse projects.
- Hands on working with the reporting tool Tableau, creating dashboards attractive dashboards and worksheets.
- Extensive work experience with Java/J2EE technologies such as Servlets, JSP, EJB, JDBC, JSF, Struts, spring, SOA, AJAX, XML/XSL, Web Services (REST, SOAP), UML, Design Patterns and XML Schemas.
- Strong experience in design and development of relational database concepts with multiple RDBMS databases including Oracle10g, MySQL, MS SQL Server & PL/SQL.
- Experience in JAVA, J2EE, WEB SERVICES, SOAP, HTML and XML related technologies.
- Have closely worked with the technical teams, business teams and product owners.
- Strong analytical and problem-solving skills and ability to follow through with projects from inception to completion.
- Ability to work effectively in cross-functional team environments, excellent communication and interpersonal skills.
TOOLS AND TECHNOLOGIES:
Hadoop/BigData Technologies: HDFS, Map Reduce, Sqoop, Flume, Pig, Hive, Oozie, Impala, Zookeeper, Ambary, Storm, Spark and Kafka
No SQL Database: HBase, Cassandra, MongoDB
Monitoring and Reporting: Tableau, Custom Shell Scripts
Hadoop Distribution: Horton Works, Cloudera, MapR
Build Tools: Maven, SQL Developer
Programming and Scripting: Java, C, C++, C# on .net, JavaScript, Shell Scripting, Python, Scala, Pig Latin, HiveQL
Java Technologies: Servlets, JavaBeans, JDBC, Spring, Hibernate, SOAP/REST services
Databases: Oracle, MY SQL, MS SQL server, Vertica, Teradata
Analytics Tools: Tableau, Microsoft SSIS, SSAS and SSRS
Web Dev. Technologies: HTML, XML, JSON, CSS, JQUERY, JavaScript
IDE Dev. Tools: Eclipse 3.5, Net Beans, My Eclipse, Oracle, JDeveloper 10.1.3, SOAP UI, Ant, Maven, RAD
Operating Systems: Linux, Unix, Windows 8, Windows 7, Windows Server 2008/2003
Hadoop/Big Data Technologies: HDFS, Map Reduce, Sqoop, Flume, Pig, Hive, Oozie, Impala, Zookeeper, Ambary, Storm, Spark and Kafka, Apache Nifi
Network protocols: TCP/IP, UDP, HTTP, DNS, DHCP
PROFESSIONAL EXPERIENCE:
Confidential, Denver, CO
Sr. Hadoop Developer
Responsibilities:
- Worked on loading disparate data sets coming from various sources to BDpaas (HADOOP) environment using SQOOP.
- Developed UNIX scripts in creating Batch load and driver code for bringing huge amount of data from Relational databases to BIGDATA platform.
- Ingested data from one tenant to the other. Developed Pig queries to load data to HBase
- Leveraged Hive queries to create ORC tables
- Created ORC tables to improve the performance for the reporting purposes. Involved in the coding and integration of several business-critical modules of CARE application using Java, spring, Hibernate and REST web services on Web Sphere application server.
- Involved in project to provide eligibility, structure and transactional feeds to River Valley Facets platform where heritage and neighborhood health plans and related commercial products are maintained and administered.
- Developed web pages using JSPs and JSTL to help end user make online submission of rebates. Also used XML Beans for data mapping of XML into Java Objects.
- Worked with Systems Analyst and business users to understand requirements for feed generation.
- Created Health Allies Eligibility and Health Allies Transactional feeds extracts using Hive, HBase, Python and UNIX to migrate feed generation from a mainframe application called CES (Consolidated Eligibility Systems) to big data.
- Used bucketing concepts in Hive to improve performance of HQL queries.
- Used numerous user defined functions in hive to attain complex business logic in feed generation.
- Developed Spark scripts by using Scala shell commands.
- Created reusable Python script and added it to distributed cache in Hive to generate fixed width data files using an offset file.
- Created a MapReduce program which looks into data in HBase current and prior versions to identify transactional updates. These updates are loaded into Hive external tables which are in turn referred by Hive scripts in transactional feeds generation.
- Worked on agile methodology using Rally
Environment: MAPR, Sqoop, Hive, Pig, Python, UNIX, HBase, Spark, Rally.
Confidential, Waukegan, IL
Sr Hadoop Developer
Responsibilities:
- Worked on different file formats like Sequence, XML, JSON files and Map files using Map Reduce Programs.
- Took important decisions of how much the poll time should be for the stream processing, what type of Hadoop stack component to use for better performance.
- Proposed and implemented a solution for their long-time issue of ordering the data in Kafka queues.
- Designed and implemented an ETL framework with the help of sqoop, pig and hive to be able to automate the process of frequently bringing in data from the source and make it available for consumption.
- Worked on importing and exporting data into HDFS and Hive using Sqoop, built analytics on Hive tables using Hive Context.
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
- Load and transform large sets of semi-structured and unstructured data on HBase and Hive.
- Implemented Map-Reduce programming model with XML, JSON, and CSV file formats. Made use of SERDE jars to load json and xml format data onto Hive tables coming from Kafka queues.
- Implemented UDFs for Hive extending Generic UDF, UDTF and UDAF base classes to change the time zones implement logic actions and extract required parameters according to the business specification.
- Extensive working knowledge of Partitioning , UDFs , Performance tuning, Compression -related properties on Hive tables.
- Developed the UNIX shell scripts for creating the reports from Hive data.
- Implemented Spark scripts in Python to perform extraction of required data from the data sets and storing it on HDFS.
- Developed spark scripts and python functions that involve performing transformations and actions on data sets.
- Configuring Spark Streaming in Python to receive real time data from the Kafka and store it onto HDFS.
- Experienced in building analytics on top of spark using machine learningSpark.ml.
- Involved in optimizing the Hive queries using Map-side join, Partitioning, Bucketing and Indexing.
- Involved in tuning the Spark modules with various memory and resource allocation parameters, setting right Batch Interval time and varying the number of executors to meet the increasing load overtime.
- Continuously monitored and managed the Hadoop cluster using Cloudera Manager.
- Used Hue for UI based PIG script execution, Tidal scheduling and creating tables in Hive.
- Created Pig Latin scripts to sort, group, join and filter the enterprise wise data.
- Involved in planning process of iterations under the Agile Scrum methodology.
- Extensive knowledge of working on Apache NiFi, used and configured of different processors to pre-process, make the incoming data uniform and format according to the requirement.
- Implemented unit testing in Java for pig and hive applications.
- Hands on experience in AWS Cloud in various AWS services such as Red shift cluster, Route 53domain configuration.
- Extensively used UNIX for shell Scripting and pulling the Logs from the Server and monitor it.
- Worked with the Data Science team to gather requirements for various data mining projects.
Environment: Cloudera CDH 5.7, Apache Hadoop 2.6.0 (Yarn), Spark 2.1.0,Spark.ml, Flume 1.7.0, Eclipse, Map Reduce, Hive 1.2.2, Pig Latin 0.17.0, Java, SQL, Sqoop 1.4.6, Centos, Zookeeper 3.5.0 and NOSQL database, Apache Nifi, AWS, S3, EMR, Red Shift Cluster.
Confidential, NYC, NY
Hadoop Developer
Responsibilities:
- Worked on extracting data from Oracle database and load to Hive database.
- Used Spark-Streaming APIs to perform necessary transformations and actions on the fly from Kafka queues in real time and persist on Cassandra using the required connectors and drivers.
- Integrated Kafka, Spark and Cassandra for streamline analytics for creating a predictive model.
- Developed Scala scripts, UDFs using both Data frames in Spark for Data Aggregation, queries and writing data back into OLTP system through Sqoop.
- Worked on modifying and executing the UNIX shell scripts files for processing data and loading to HDFS.
- Worked extensively on optimizing transformations for better performance.
- Was involved in carrying out the important design decisions in creating UDFs, partitioning the data in hive tables at two different levels based on the related columns for efficient retrieval and processing of queries.
- Tweaked lot of options to get performance boost like trying it out with different executer count and memory options.
- My team was also involved in maintenance, adding the feature of stable time zones across all records in the database.
- Uploaded and processed more than 30 terabytes of data from various structured and unstructured, heterogeneous sources into HDFS file system using Sqoop and Flume enforcing and maintaining the uniformity across all the tables.
- Developed complex transformations using HiveQL to build aggregate/summary tables.
- Used Solr/Lucene for indexing and querying the JSON formatted data.
- Implemented unit tests for pig and hive applications.
- Developed UDF's in Python to implement functions according to the specifications.
- Developed Spark scripts, configured according to business logic, good knowledge of actions available.
- Well versed with the HL7 international standards as the data was organized according to this format.
- Formatted and built analytics on top of the data sets that were complied with HL7 standards.
- Created UDFs in scala, java for formatting and applying transformations on the information in HL7 versions.
- Analyze the JSON data using hive SerDe API to deserialize and convert into readable format.
- Involved in increasing and optimizing the performance of the application using Partitioning and Bucketing on Hive tables, developing efficient queries by using Map-side joins and Indexes.
- Worked with downstream team in generating the reports on Tableau.
- Conducted code reviews to ensure systems operations.
Environment: CDH 5.1.x, Hadoop 2.2.0, HDFS, Map Reduce, Sqoop, Flume, Hive 2.0.x, SQL Server, TOAD, Oracle, Scala 2.9.1,Solr/Lucene, PL/SQL, Eclipse, JAVA, Shell scripting, Vertica, Unix, Cassandra, HL7 standard.
Confidential, Bloomington, IL
Hadoop Developer
Responsibilities:
- Involved in architecture design, development and implementation of Hadoop deployment, backup and recovery systems.
- Developed MapReduce programs in Python using Hadoop streaming API to parse the raw data, populate staging tables and store there fined data in partitioned HIVE tables.
- Enabled speedy reviews and first mover advantages by using Oozie to automate data loading into the Hadoop Distributed File System and Pig to pre-process the data.
- Converted applications which was on map-reduce architecture to Spark using Python API which performed the business logic.
- Involved in creating Hive tables, loading with data, writing hive queries that will run internally in map reduce way.
- Imported Teradata datasets onto the HIVE platform using Teradata JDBC connectors.
- Was involved in writing FastLoad and MultiLoad scripts to load the tables
- Worked with diverse types of Indexes and Collect Statistics in Teradata and improving of execution strategy.
- Worked with the SQL assistant and BTEQ to ingest and execute queries, stored procedures and update the tables.
- Worked in extracting XML type files using XPath and storing it onto to Hive tables.
- Developed multiple Kafka Producers and Consumers as per the software requirement specifications.
- Involved in designing the tables in Teradata while importing the data.
- Developed the UNIX shell scripts for creating the reports from Hive data.
- Experienced in managing and reviewing the Hadoop log files.
- Developed Hive jobs to parse the logs, structure them in tabular format to facilitate effective querying on the log data.
- Extensively used UNIX for shell Scripting and pulling the Logs from the Server.
- Worked on different file formats like Sequence files, XML files and Map files using Map Reduce Programs.
- Worked with Avro Data Serialization system to work with JSON data formats.
- Implemented the workflows using Apache Oozie framework to automate tasks.
- Completed testing of integration and tracked and solved defects.
- Used AWS services like EC2 and S3 for small data sets.
- Involved in loading data from UNIX file system to HDFS
Environment: Hadoop Hortonworks 2.2, Hive, Pig, HBASE, Sqoop and Flume, Oozie, AWS, S3, EC2, EMR Spring, Kafka, SQL Assistant, Python, UNIX, Teradata
Confidential, Dyersburg, TN
Jr. Hadoop Developer
Responsibilities:
- Supported and monitored Map Reduce Programs running on the cluster.
- Evaluated business requirements and prepared detailed specifications that follow project guidelines to develop programs.
- Configured the Hadoop cluster with Name-node, Job-tracker, and Task-trackers on slave nodes and formatted HDFS.
- Used Oozie workflow engine to schedule and execute multiple Hive, Pig and Spark jobs by passing arguments.
- Involved in creating Hive Tables, loading data and writing Hive queries to invoke and run Map Reduce jobs in the backend.
- Designed and implemented Incremental Imports into Hive tables using Sqoop and cleaning the staging tables and files.
- Involved in collecting, pre-processing, aggregating and moving data from servers to HDFS using Apache Flume.
- Developed multiple Map Reduce jobs in python for data cleaning and preprocessing.
- Exported the result set from Hive to MySQL using Sqoop after processing the data.
- Analyzed the data by performing Hive queries and running Pig scripts to study customer behavior.
- Have hands on experience working on Sequence files, AVRO, HAR file formats and compression.
- Used Partitioning and Bucketing in Hive tables to increase the performance and parallelism.
- Experience in writing Map Reduce programs in python to cleanse Structured and unstructured data.
- Wrote Pig Scripts to perform ETL procedures on the data in HDFS.
- Implemented Hive and Pig scripts to analyze large data sets.
- Loaded and transformed large sets of structured, semi structured and unstructured data onto HBase column families.
- Created HBase tables to store data coming from different portfolios and created Bloom filters on column families for efficiency.
- Worked on improving the performance of existing Pig and Hive Queries by optimization.
- Analyzed the partitioned and bucketed data and compute various metrics for reporting.
- Deployed and worked with Apache Solr search engine server to help speed up the search of the sales and production data.
- Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the data pertaining to different states across northern-USA.
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig Latin scripts and store these intermediate results for further analytics.
- Migrated ETL jobs to Pig scripts to perform the Transformations, joins and pre-aggregations before storing data onto HDFS.
- Involved in loading data from RDBMS and web logs into HDFS using Sqoop and Flume.
- Worked on loading the data from MySQL to HBase where necessary using Sqoop
- Co-ordinate with offshore and onsite team to understand the requirements to propagate and prepare design documents from the requirements specification and architectural designs.
- Worked with application teams to install Operating Systems, Hadoop updates, patches, and version upgrades as required.
Environment: Hadoop1x, Hive, Pig, HBASE, Sqoop and Flume, Spring, jQuery, Java, J2EE, HTML, JavaScript, Hibernate, PL/SQL, Windows NT, UNIX Shell Scripting, Putty and Eclipse.
Confidential
Java developer
Responsibilities:
- Implemented Microsoft Visio and Rational Rose for designing the Use Case Diagrams, Class models, Sequence diagrams, and Activity diagrams for SDLC process of the application.
- Deployed GUI pages by using JSP, JSTL, HTML, DHTML, XHTML, CSS, JavaScript,AJAX
- Configured the project on WebSphere 6.1 application servers
- Implemented the online application by using Core Java, Jdbc, JSP, Servlets and EJB 1.1, WebServices, SOAP,WSDL
- Used Log4J logging framework to write Log messages with various levels.
- Involved in fixing bugs and minor enhancements for the front-end modules.
- Performed live demos of functional and technical reviews.
- Maintenance in the testing team for System testing/Integration/UAT.
- Guaranteeing quality in the deliverables to the product owners and business team.
- Conducted Design reviews and Technical reviews with other project stakeholders.
- Implemented Action Classes and Server-side validations for account activity, registration and Transaction’s history.
- Designed user-friendly GUI interface and Web pages using HTML, CSS, Struts, JSP.
- Involved in writing Client-Side Scripts using Java Scripts and Server-Side scripts using Java Beans.
Environment: JDK 1.5, JSP, WebSphere, JDBC, EJB2.0, XML, DOM, SAX, XSLT, CSS, HTML, JNDI, WEB SERVICES, WSDL, SOAP, RAD, PL/SQL, JavaScript, DHTML, XHTML, JavaMail,PL/SQL DEVELOPER, TOAD, POI REPORTS, WINDOWS XP, RED HAT LINUX
Confidential
Web Developer
Responsibilities:
- Involved in various stages of Enhancements in the Application by doing the required analysis, development, and testing.
- Prepared the High and Low-level design document and Generating Digital Signature
- For analysis and design of application created Use Cases, Class and Sequence Diagrams.
- For the registration and validation of the enrolling customer developed logic and code.
- Developed web-based user interfaces using struts framework.
- Coded and deployed JDBC connectivity in the Servlets to access the Oracle database tables on Tomcat web-server.
- Handled Client-side Validations used JavaScript
- Involved in integration of various Struts actions in the framework.
- Used Validation Framework for Server-side Validations
- Created test cases for the Unit and Integration testing.
- Front-end was integrated with Oracle database using JDBC API through JDBC-ODBC Bridge driver at server side.
Environment: Java Servlets, JSP, Java Script, XML, HTML, UML, Apache Tomcat, JDBC, Oracle, SQL, Log4j
