Sr. Big Data Architect Resume
Troy, NY
SUMMARY:
- Above 10+ years of professional IT experience in all phases of Software Development Life Cycle which includes hands on experience in Java Technologies and Big Data.
- Experienced in ingestion, storage, querying, processing and analysis of Big Data.
- Expertise in Micro services automation for Micro service oriented Enterprise platform.
- Worked extensively writing queries in Hive/Impala to verify and explore data.
- Experienced in Hadoop Ecosystem development using MapReduce, HDFS, Hive, Pig, HBase, Sqoop, Flume, Spark, and Oozie.
- Strong knowledge on Hadoop architecture and various components such as HDFS, JobTracker, Task Tracker, NameNode, DataNode and MapReduce programming paradigm including Software Architecture, Object oriented programming, Big Data.
- Worked in environments using Agile(Scrum), RUP and Test Driven development methodologies.
- Hand on experienced of multiple distributions Cloudera, Hortonworks.
- Hands on experienced in writing MapReduce programs using Java to handle different data sets using Map and Reduce tasks.
- Designed HIVE queries &Pig scripts to perform data analysis, data transfer and table design.
- Expertise in loading and transforming large sets of structured, semi - structured and unstructured data.
- Experienced with cloud: Hadoop-on-Azure, AWS/EMR, Cloudera Manager (also direct-Hadoop-EC2 (non EMR)).
- Hands on experienced in designing and creating Hive tables using shared meta-store with partitioning and bucketing.
- Developed Spark scripts by using Scala shell commands as per the requirement.
- Experienced in performance tuning the Hadoop cluster by gathering and analyzing the existing infrastructure.
- Experience in writing shell scripts to move the Shared data from MySQL servers to HDFS.
- Written multiple MapReduce programs in Python for data extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV and other compressed file formats.
- Expertise in designing and coding Stored Procedures, Triggers, Cursers and Functions using PL/SQL.
- Experienced in writing Hive Queries to generate reports using Hive Query Language.
- Experience in managing and reviewing Hadoop log files.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python.
- Experienced with Operating Systems like Windows, Linux, and Macintosh.
- Involved extensively in designing/developing web based applications using HTML and MVC design patterns.
- Data Warehousing experienced using TeradataV2r5/V12, BTEQ, Tpump, Fastload, Multiload and FastExport.
- Extensive experience in Tomcat Server, JBoss, Weblogic and WebSphere application
- Good Experience in implementing Web Services such as SOAP, WSDL, UDDI, REST.
- Proficient in using RDBMS concepts with Oracle10g/12c, MySQL and experienced in writing SQL, PL/SQL Stored procedures.
- Has the ability to provide solutions from a functional and technical perspective, meet deadlines.
- Team player, learning skills, technically competent and result-oriented with problem solving skills.
TECHNICAL SKILLS:
HADOOP ecosystem: Hadoop and MapReduce, Sqoop, Hive, PIG, HBASE, YARN, HDFS, Zookeeper, Oozie, Flume, Spark, Kafka
Languages: Java, HTML, DHTML, XML, CSS, UNIX Shell Script, JavaScript, SQL, PL/SQL
Technologies: JSP, Servlets, JNDI, JDBC, EJB, JMS, Java Beans, SOAP, REST AJAX, AWT, Swings
Distributed Technologies: RMI, EJB, JMS
Application Server: JBoss, Apache Tomcat 5.5/6.0
J2EE Frameworks: Struts 2.0, ANT build tool, Log4J, MVC, Hibernate
IDE s: Eclipse, JBuilder, NetBeans
Database(s): Oracle(10g,12c), DB2, MySQL
Version Control Tools: Rational Clear Case, GitHub, Perforce, SVN
Case Tools: Rational Rose, UML, OOAD
Windows Family, MS: DOS, UNIX, Linux, Macintosh.
PROFESSIONAL EXPERIENCE:
Confidential, Troy, NY
Sr. Big Data Architect
Responsibilities:
- Involved in Design and Architecting of Big Data solutions using Hadoop Eco System.
- Collaborate in identifying the current problems, constraints and root causes with data sets to identify the descriptive and predictive solution with support of the Hadoop HDFS, MapReduce, Pig, Hive, and Hbase and further to develop reports in Tableau.
- Involved in analysis, design and development phases of the project. Adopted Agile methodology throughout all the phases of the application.
- Worked on analyzing Hadoop cluster using different Bigdata analytic tools including Kafka, Sqoop, Storm, Spark, Pig, Hive and Map Reduce.
- Installed/Configured/Maintained Hortonworks Hadoop clusters for application development and Hadoop tools like Hive, Pig, HBase, Zookeeper and Sqoop.
- Architect the Hadoop cluster in Pseudo distributed Mode working with Zookeeper and Apache.
- Storing and loading the data from HDFS to AmazonAWSS3 and backing up and Created tables in AWS cluster with S3 storage.
- Utilized Big Data technologies for producing technical designs, prepared architectures and blue prints for Big Data implementation and involved in to writing Scala program using sparkcontext.
- Migrated physical data center environment to AWS also designed, built, and deployed a multitude applications utilizing almost all of the AWS stack (EC2, S3, RDS, )
- Provided technical assistance for configuration, administration and monitoring of Hadoop clusters.
- Involved in loading data from Linux file system to HDFS and Importing and exporting data into HDFS and Hive using Sqoop and Kafka.
- Implemented Partitioning, Dynamic Partitions, Buckets in Hive and Supported MapReduce Programs those are running on the cluster.
- Prepared presentations of solutions to BigData/Hadoop business cases and present the same to company directors to get go-ahead on implementation.
- Successfully integrated Hive tables and Mongo DB collections and developed web service that queries Mongo DB collection and gives required data to web UI.
- Installed and configured Hadoop Map Reduce, HDFS, Developed multiple Map Reduce jobs in Java for data cleaning and preprocessing.
- Used Spark to create API's in Java and Scala and real time streaming the data using Spark with Kafka and developed Hive queries, Pig scripts, and Spark SQL queries to analyze large datasets.
- Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scala.
- Implemented Storm builder topologies to perform cleansing operations before moving data into Cassandra and transfer data between databases using Sqoop.
- Worked on debugging, performance tuning of Hive & Pig Jobs and implemented test scripts to support test driven development and continuous integration.
- Developed enhancements to MongoDB architecture to improve performance and scalability.
- Deployed Algorithms in Scala with Spark, using sample datasets and done Spark based development with Scala.
- Manipulating, cleansing & processing source data and stage it on final hive/redshift tables and involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs.
- Used Storm to consume events coming through Kafka and generate sessions and publish them back to Kafka.
- Gathered and analyzed the requirements and designed class diagrams, sequence diagrams using UML.
- Writing scala classes to interact with the database and writing scala test cases to test scala written code.
- Performed exceptional J2EE Software Development Life Cycle (SDLC) of the application in Web and client-server environment using J2EE.
- Used Kibana web - based data analysis and dash boarding tool for elastic search and used logstash to stream data from one or many inputs, transforms it and output it one or many outputs.
Environemnt: Big Data, Hadoop, HDFS, Pig, Hive, MapReduce, Azure, Sqoop, Spark, Kafka, LINUX, Cassandra, MongoDB, Scala, Storm, Elastic search, SQL, PL/SQL, Scala, AWS, S3, Informatica, Redshift.
Confidential, Phoenix, AZ
Big Data/Hadoop Developer
Responsibilities:
- Worked on analyzing Hadoop cluster and different big data analytic tools including Pig, HBase database and Sqoop.
- Data Ingestion into the Indie-Data Lake using Open source Hadoop distribution to process Structured, Semi-Structured and Unstructured datasets using Open source Apache tools like FLUME and SQOOP into HIVE environment.
- Upgraded the Hadoop Cluster from CDH3 to CDH4, setting up High Availability Cluster and integrating HIVE with existing applications.
- Develop predictive analytic using Apache Spark Scala APIs.
- Developed MapReduce jobs in Java API to parse the raw data and store the refined data.
- Configured Sqoop and developed scripts to extract data from MYSQL into HDFS.
- Created tables in HBase to store variable data formats of PII data coming from different portfolios.
- Implemented best income logic using Pig scripts.
- Involved in identifying job dependencies to design workflow for Oozie & YARN resource management.
- Working on data using Sqoop from HDFS to Relational Database Systems and vice-versa. Maintaining and troubleshooting
- Exploring with Spark to improve the performance and optimization of the existing algorithms in Hadoop using Spark context, Spark-SQL, Data Frame, pair RDD's.
- Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
- Used Cloudera Manager for installation and management of Hadoop Cluster.
- Involved in scheduling Oozie workflow to automatically update the firewall.
- Developing data pipeline using Flume, Sqoop, Pig and Java map reduce to ingest customer behavioral data and financial histories into HDFS for analysis.
- Involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs.
- Worked in the BI team in the area of Big Data Hadoop cluster implementation and data integration in developing large-scale system software.
- Worked on MongoDB, HBase (NoSQL) databases which differ from classic relational databases
- Worked in AWS EC2, configuring the servers for Auto scaling and Elastic load balancing
- Implemented test scripts to support test driven development and continuous integration.
- Developed Spark code using Scala and Spark-SQL for faster testing and data processing.
- Responsible for managing data coming from different sources
- Imported millions of structured data from relational databases using Sqoop import to process using Spark and stored the data into HDFS in CSV format.
- Worked in tuning Hive & Pig to improve performance and solved performance issues in both scripts
- Used Spark SQL to process the huge amount of structured data.
- Developed Spark streaming application to pull data from cloud to Hive table.
Environment: HDFS, Map Reduce, Pig, Hive, Sqoop, Flume, Oozie, HBase, Impala, Spark Streaming, Yarn, Eclipse, spring, PL/SQL, UNIX Shell Scripting, Cloudera.
Confidential, Cincinnati, OH
Big Data/Hadoop Developer
Responsibilities:
- Involved in Requirement Analysis, Development and Documentation.
- Developing Scripts to schedule various Hive, Pig and Sqoop Jobs.
- Developed Map-Reduce programs to clean and aggregate the data
- Experienced in Importing and exporting data into HDFS and Hive using Sqoop.
- Should be able to define open source based hybrid framework for Micro services Automation.
- Experience in Hadoop development using HDFS, Map Reduce, Pig, Hive, Impala Sqoop.
- Used Talend reusable components context variable and global Map variables
- Designed and implemented a data ingesting using flume into HDFS.
- Installed Hadoop, MapReduce, HDFS, AWS and developed multiple MapReduce jobs in PIG and Hive for data cleaning and pre-processing.
- Full stack software developer primarily using the following technologies: MongoDB, ExpressJs, AngularJS, NodeJs (MEAN Stack).
- Load the data into Spark RDD and do in memory data Computation to generate the Output response.
- Developed a data pipeline using Kafka and Storm to store data into HDFS.
- Worked on Apache Spark along with SCALA Programming language for transferring the data in much faster and efficient way.
- Experienced in Ingesting real time into HBASE using Kafka through Spark Streaming
- Writing Pig Latin scripts to process the data and written UDF in Java and Python.
- Developed Spark Streaming applications for Real Time Processing.
- Worked on importing and exporting data from Oracle and DB2 into HDFS and HIVE using Sqoop.
- Configure Oozie workflow to run multiple Hive and Pig jobs which run independently with time and data availability.
- Used Flume to collect, aggregate, and store the web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
- Optimized the Hive tables using optimization techniques like partitions and bucketing to provide better performance with HiveQL queries.
- Developed HBase data model on top of HDFS data to perform real time analytics using Java AP
- Implemented Fair schedulers on the Job tracker to share the resources of the Cluster for the Map Reduce jobs given by the users.
- Worked extensively on development and maintenance of HADOOP applications using JAVA and MapReduce.
- Used AWS to produce comprehensive architecture strategy for environment mapping.
- Configured Oozie workflow to run multiple Hive and Pig jobs which run independently with time and data availability.
- Used Oozie scheduler system to automate the pipeline workflow.
- Involved in managing and reviewing Hadoop log files.
- Experienced in handling data from different data sets, join them and pre-process using Pig join operations.
Environment: Hadoop, HDFS, Hive, HBase, Spark, Kafka, Java (JDK 1.7), Linux, XML, MySQL, MySQL Workbench, Eclipse, PL/SQL, Sub Version, Windows NT, UNIX Shell Scripting, Map Reduce, Python, Flume, Pig, Scala, Yarn, Sqoop, Zoo Keeper, Cloudera, Oozie, Cassandra, NoSQL, Talend Big Data studio, ETL, agile, Teradata.
Confidential, Minneapolis, MN
Big Data Engineer
Responsibilities:
- Developed MapReduce programs in Java for parsing the raw data and populating staging tables.
- Stack app using MongoDB, Node.js + Express, JavaScript, jQuery
- Created Talend Mappings to populate the data into Staging, Dimension and Fact tables.
- Expertise in Micro services test scripting using open source tools (ReST/API services automation and E2E functions scripting at services layer)
- Deployed Hadoop Cluster in the following nodes, Managing and scheduling jobs on Hadoop cluster.
- Preparation of high-level design for the new requirements and Development of the code using Hadoop eco systems Pig, Hive, Impala as per the business requirements
- Implemented Generic writable to in corporate multiple data sources into reducer to implement recommendation based reports using MapReduce programs.
- Analyzed structured data using HIVE tool.
- Proficiency and knowledge in understanding the architectural components of Hadoop like HDFS, JobTracker, TaskTracker, NameNode, DataNode and MapReduce concepts.
- Involved in creating Hive tables, loading data and writing hive queries that will run internally in mapreduce way.
- Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from a variety of data sources.
- Worked on SCALA Programming language which is supported by APACHESPARK
- Enabled speedy reviews and first mover advantages by defining the job flow in Oozie to automate data loading into the Hadoop Distributed File System and PIG to pre-process the data.
- Using Sqoop, exported data from HDFS to SQL Server for provisioning.
- Designed and executed various simple enrichments to the data using Pig and Java
- Performed partitioning on monthly bases.
- Worked on HBase. Configured MySQL Database to store Hive metadata.
- Involved in loading data from UNIX file system to HDFS.
- Involved in creating Hive tables, loading with data and writing hive queries.
- Responsible for performing extensive data validation using Hive.
- Performed various operations on data lake of MapR cluster which involves moving the data, enriching the data, performing validations.
- Automated all the jobs, for pulling data from FTP server to load data into Hive tables, using Oozie workflows.
- Experience in managing and reviewing Hadoop log files.
- Experience in build scripts using Maven and do continuous integrations systems like Jenkins.
- Used JIRA for bug tracking.
- Used Tableau for visualizing and to generate reports.
Environment: Java, Hadoop, Linux, MapReduce, HDFS, Hive, Shell Scripting, Java (JDK 1.6), Eclipse, SVN, JIRA, Cassandra, Shell Scripting, MySQL, DB Visualizer, Linux, Sqoop, Apache Hive, Apache Pig
Confidential, Boston, MA
Java/J2EE Developer
Responsibilities:
- Worked on Full Cycle of Software Development from Analysis through Design, Development, Testing, Integration, Deployment.
- Designed and developed User Interface of application modules using HTML, JSP, CSS, JavaScript, jQuery and AJAX.
- Involved in complete requirement analysis, design, coding and testing phases of the project.
- Extensively worked on Talend Designer Components - Data Quality (DQ), Data Integration (DI) and Master Data Management (MDM)
- Implemented the project according to the Software Development Life Cycle (SDLC).
- Developed JavaScript behavior code for user interaction.
- Used HTML, JavaScript, and JSP and developed UI.
- Worked with spring framework.
- Used JavaScript, JQuery, AJAX, JSTL, CSS and Struts2 tags for developing the JSP'S.
- Developed client side screen using JSP, HTML and DHTML.
- Used Eclipse as IDE.
- Used Hibernate for connecting to the database and mapping the entities by using hibernate annotations. implementation.
- Wrote JUNIT test cases for testing all spring service calls and Spring MVC validations.
- Worked on generating the Web Services classes, WSDL using Apache Axis.
- JSON is used for serializing and de-serializing data that is sent to or receives from JSP pages.
- Creating environment for user-acceptance testing and facilitating Integration and User Acceptance Testing with JUNIT
- Used GIT as version control system to manage the progress of the project.
- Developed POJOs and Java beans to implement business logic.
- Developed presentation layer using Java Server Faces (JSF) MVC framework.
- Used JSP, JSTL, HTML and CSS, JQuery as view components in MVC.
- Coordinate testing meetings (e.g. status update; action items; open issues; prioritizing errors; Communicate Priorities)
- Ensure all open issues/and or risks are documented prior to moving to next testing stage.
- Designed database and created tables, written the complex SQL Queries and stored procedures as per the requirements.
Environment: J2EE, Java Script, JSP, Servlets, Win Runner, GUI, Test Director 8.0, DB2,, IBM Mainframe, JBOSS, JSF, JSTL, Java Script, HTML, JQuery, CSS, JUnit, Web Sphere, UML, JSON, XML, Web Services, WSDL, Apache Maven, UNIX, Windows' XP Professional.
Confidential, Minneapolis, MN
Java/J2EE Developer
Responsibilities:
- Involved in designing and developing the application using JSP, HTML, CSS and made client validations using JavaScript
- Understanding the business requirements and preparing the design document.
- Developed user interface using JSP, Struts Tag Libraries to simplify the complexities of the application.
- Created REST based web service in JSON, RSS and CSV format.
- Created java classes to communicate with database using JDBC.
- Developed Web Services using SOAP, WSDL and JAX-WS programming model
- Hibernate is used as persistent at middle tire for providing object model over relational data model
- Developed data access objects to interact with persistent layer using hibernate session for CRUD operations.
- Designed Java Servlets and Objects using J2EE standards.
- Experience in JSP, JSF, Servlet, JAVA Bean, spring, Webservices, and Spring Boot development.
- Designed and developed the application using various design patterns, such as session facade, business delegate and service locator.
- Using Linux Command and Shell scripts to create deployment script, check Linux and JBoss Server performance and Error log.
- Written Approach Notes documents for all NM upgrades and Client Specific Re-Branding activities
- Created new tables, Sequences and written SQL queries and PL/SQL in Oracle.
- Helping manager in handling risk assessment and subsequently created contingency and mitigation plans
- Used JSP, JavaScript, JSTL, EL, Custom Tag libraries, Tiles and Validations provided by struts framework.
- Used SVN as version control system for the source code.
- Used JAXP (DOM, XSLT), XSD for XML data generation and presentation.
- Wrote Junit test classes for the services and prepared documentation.
Environment: Java, SOA, JMS, Angular JS, JMX, IBM MQ Series, Node JS, Web Services, Axis, SOAP UI, Hibernate, JNDI, XML, XSD, JAXB, JAXP, JDBC, bootstrap, Spring, Junit, JProfile, Ant, JPA, JTA, JDBC, Maven, PL/SQL Developer, DB2, Unix, Log4J, UML and Agile, IBM RAD.
Confidential
Java/J2EE Developer
Responsibilities:
- Enhanced the existing systems using Java, J2EE, JSP, Spring, Hibernate, Restful, Web Services and Java Beans
- Enhanced Cross connections vs path finding process using java multi-threading techniques for DN2, and DB2 nodes
- Used JDBC to connect to the backend database and developed stored procedures.
- Developed SQL statements for updating and accessing data from database.
- Used SVN for version control across common source code used by developers.
- Developed the user interface using JSP and Java Script to view all online trading transactions.
- Developed both Session and Entity beans representing different types of business logic abstractions. Used application server like WebSphere and Tomcat.
- Used Rational Application Developer (RAD) for developing the application.
- Used Spring DAO to connect t the database.
- Designed Java Servlets and Objects using J2EE standards.
- Implemented the application using the concrete principles laid down by several design patterns such as MVC, Business Delegate, Data Access Object, and Singleton.
- Produced and Consumed REST web services using Apache CXF and spring.
- Used Maven and configured Jenkins to build and deploy the application.
- Used JUnit framework for unit testing and ANT to build and deploy the application on WebLogic Server.
- Used Ant builds script to create EAR files and deployed the application in Web Logic server.
- Performed Smoke testing for various builds before accepting into the QA environments.
- Confidential version control system has been used to check-in and checkout the developed artifacts. The version control system has been integrated with Eclipse IDE.
Environment: Java, Oracle 8i/10g, Eclipse Luna, JBoss, Spring MVC, Junit, JMockit, Web services,, Java/J2EE, SQL, PL/SQL, JSP, EJB, Struts, Hibernate, WebLogic, HTML, AJAX, Java Script, JDBC, XML, JMS, XSLT, UML, JUnit, log4j, MyEclipse, Star UML, SVN.
