Big Data Developer /spark Developer Resume
Bloomington, IL
SUMMARY
- Over 12+ years of Experience in Big Data/Hadoop Echo systems, AWS Certified Solution Architect that includes Development and Maintenance of various programs and applications.
- Veritable experience working with world class organizations viz. Group.1 Insurance Company, Confidential, DXC Technology, Confidential Data and NTT Data Americas.
- Business and Functional Knowledge in various domains like Insurance and Transportation (Includes Rail Road and Airline).
- Developed and Maintained many applications/Programs in Security and HCM HRMS Project for Group 1 Insurance Company, Corporate Systems, Corporate Loyalty Programs for Airlines and Electronic Data Interchange (EDI), Service Management Systems, TAXI, Freight Claims, Drug and Alcohol Random Testing(DARTS) for Rail Road.
- Having good knowledge and Hands on experience on BIG DATA Analytics implementation using Hadoop echo systems like HDFS, HBASE, MAP REDUCE, SQOOP, Apache HIVE, Apache Pig, Apache OOZIE, Apache Flume, Apache Spark, Scala and High Level Language Python (NumPy, SciPy, scikit - learn, pandas).
- Experience in managing Hadoop clusters usingCloudera Manager (CM) and Experience in managing and reviewingHadoop log files.
- Extensive experience working inOracle, DB2, SQLServer andMy SQLdatabase.
- Developed scalable and reliable data solutions to move data across systems from multiple sources in real time as well as batch modes.
- Excellent understanding / knowledge ofHadoop architecture and Map Reduce programming, and Involved in writing the Pig scripts to reduce the job execution time. Proficient in Installation, Configuration of Hadoop and its echo systems.
- Knowledge in implementing advanced procedures like text analytics and processing using Apache Spark with Python language.
- Experience in designing and handling of various Data Ingestion patterns (Batch and Near Real Time) using Sqoop, Distcp, Apache Storm, Flume and Apache Kafka.
- Experience in designing and handling of various Data Transformation/Filtration patterns using Pig, Hive, Python.
- Strong Knowledge on Hadoop architecture and various components such as HDFS, Job and Task Tracker, Name and Data Node, Secondary Name Node and Map Reduce programming.
- Proficient in Various design patterns of Data Ingestion, Data Modelling and In-depth understanding ofData Structureand Algorithms.
- Proficient in writing stored procedures, Complex SQL Queries, optimizing the SQL to improve performance, Packages, Functions and Database Triggers using SQL and Possess Strong data analysis skills using Python, Hive, Apache Spark, MS Excel and Access DB.
- Having good Knowledge and understanding of Amazon Web Services (AWS) with deep expertise in Amazon’s cloud computing offerings.
- Expert-level knowledge of Amazon EC2, Amazon S3, Amazon SimpleDB, Amazon RDS, Amazon Elastic Load Balancing, Amazon SQS, and other services of the AWS family
- Experience in using IDE's such as Eclipse, NetBeans for debugging and using java coding standards.
- Good working knowledge on creating/maintaining/understanding of High-Availability, Fault Tolerance, Scalability, Database Concepts, System and Software Architecture, Security, IT Infrastructure, Virtualization, and Internet Technologies
- Coordination of Production Releases, including major, minor, and hotfix releases. Also, Prepared technical design documents.
- Demonstrated ability to quickly learn new tools and paradigms to deploy cutting edge solutions.
- Proficient in mapping business requirements, use cases, scenarios, business analysis, and workflow analysis. Act as liaison between business units, technology and IT support teams.
- Good at Writing reusable, testable, and efficient code and Ability to integrate multiple data sources and databases into one system.
- Good at working as a Team player/Technical Lead and under less supervision and worked in the Onsite-Offshore model and have experience in managing remote teams. Excellent time-management and task-management skills and Self-motivated, excellent communication, analytical, interpersonal and presentation skills.
- Thorough understanding of SDLC various phases like Requirements, Analysis/Design, Development, Testing, Implementation and Maintenance and Working knowledge of project methodologies (i.e. waterfall, agile), familiarity with Systems Technical Architectures, experience & familiarity with implementation processes, procedures on various platforms, knowledge.
- Experience in preparing effort estimates, task schedules, Production builds and change requests useful for a live release.
- Experienced in creation of the documentation needed for the project implementation topics such as rollouts, contingency plans, communications, dependencies, etc. and Perform Annual BCP (Business Continuity Plan) and Disaster recovery.
- Having good understanding of the end-to-end Systems processes and the interfaces and dependencies with other processes, application design documents, functionality, data flow and technical aspects.
- Ability to multi-task for different applications and Ability to take ownership and provide coordination for tasks related to implementation to test and production.
- Ability to deliver high-quality results under tight deadlines and ability to work In Team/Multi Diverse stake holder environment.
- Good Facilitation and Coordination skills to guide and direct key parts of project work with team such as meetings, planning sessions, and training of team members.
- Expertise in developing Use Cases, Sequence Diagrams and Class Diagrams and Adhere to SCM (Software Configuration Management) during implementation.
- Having excellent coordination skills to gather information, details and people to implement a project successfully and, excellent written and verbal communication Skills.
- Possess strong interpersonal and, excellent analytical & problem solving skills.
- Ability to adapt to evolving technology strong sense of responsibility and accomplishment.
TECHNICAL SKILLS
Hadoop/Big Data Technologies: Hadoop (Cloudera): HDFS, Map Reduce, Pig, HBase, Kafka, Spark, Zookeeper, Hive, Oozie, Sqoop, Flume, Storm, Impala.
Programming Languages: Python, Scala, HTML, SQL, PL/SQL, COBOL, JCL, REXX, FOCUS, DB2 Stored Procedures, SQL, XML, Windows Batch Scripts, Java JDK1.4/1.5/1.6 (JDK 5/JDK 6), Linux shell scripting.
Operating Systems: UNIX, Windows, LINUX
Application Servers: IBM Web sphere, Web Sphere, CICS
Deployment Tools: DMS (Deployable Management System), IBM Urban Code/Jenkins, Change Management
Messaging Services: Message Broker, MQ, Apache Kafka, Amazon SQS
Databases: Netezza & MySQL 4.x/5.x, Oracle, IBM DB2, IMS DB/DC
Project Management Tools: VSS (Visual Source Safe), SharePoint, REMEDY, TRAC
Transfer Protocols: FTP, TELNET, EDIFACT
IDE’s: Eclipse 3.x, IBM Web Sphere Application Developer, IBM RAD 7.0
Tools: SVN, GitHub, TOAD, SQL Developer, Maven, Application Designer, App Engine, CI, SQR, SQL, PS-Query, People code, Data Mover, Eclipse, BMCDB2, DB2MENU, IBM DATA STUDIO, OPTIM, MAINVIEW, KADET, SIAANACONDA, GTB (CICS Map Design Tool), IBM Serena Changeman, ROVR (Registration Ownership Versioning Route), RMS, Troux, SPAR.
PROFESSIONAL EXPERIENCE
Big Data Developer /Spark Developer
Confidential, Bloomington, IL
Responsibilities:
- Teamed up with Architects to design Spark model for processing security logs and identifying Anomalies using Graph Network Analysis with Driver Node, 27 Nodes/Executors and each having 20 Virtual Cores running on YARN Application Manager.
- Designed and Developed Graph Network Analysis using Node Level and community anomalies using Egonet Method and Traffic Dispersion Method by using Scala, Python, Spark and GrapX Library.
- Designed and developed data ingestion patterns for loading data from various sources to HDFS using Flume, Kafka and SQOOP
- As part of Data Filtration, Parsing and Transforming, Written and executed Apache PIG scripts on top of the HDFS data, Created Hive tables to store the processed results in a tabular format.
- Developed scalable and reliable data solutions to move data across systems from multiple sources in real time(Flume/Kafka) as well as batch modes (SQOOP, CLI, DISTCP, MDT).
- Responsible in Managing and scheduling Jobs on a cluster using OOZIE.
- Utilized expertise in models that leverage the newest data sources, technologies, and tools, such as Python, Hadoop, Spark, AWS, as well as other cutting-edge tools and applications for Big Data.
- Developed the Sqoop scripts to make the interaction between Pig and MySQL Database.
- Used RDD's to perform transformation on datasets as well as to perform actions like count, reduce, first.
- Implemented various checkpoints on RDD's to disk to handle job failures and debugging.
- Develop integration solution to bridge the gap between external and HDFS and Reduced the model run times by performance tuning of the models for the business to run the model hundreds of times a day.
- Good knowledge on Spark platform parameters like memory, cores and executors.
- Used Spark DataFrame API over Cloudera platform to perform analytics on hive data. also, by Using Python Libraries NumPy, SciPy, scikit-learn, pandas analyzed large datasets and developed graphs.
- Implemented test scripts to support test driven development and continuous integration.
- Used version control tools like IBM Urban code/Jenkins for deployments.
- Created Control M jobs for automation for workflows and Experienced in fixing various production issues during user acceptance test.
- Prepared Project Management documents like KCD (Knowledge capture Document), Topology Diagrams, and Application related documents.
- Developed Python Programs for generating ITSS Reports and Process Scheduler Alerts and Processing of Jobs.
- With use of Analytic Languages Like Python and VBB developed tools to help HCM Technical team for their daily usage.
- Helped Application Teams to fix the data issues, Oracle delivered COBOL issues and PIA issues and Co-ordinate with Oracle for resolving of Oracle SR’s in HCM application.
- Responsible for adding and maintaining assets (Meta data) and Sensitivity in IIS (IBM Infosphere Information Server).
- Responsible to conduct Annual Access Cleanups, DSA, AUTHID Cleanups and worked with Enterprise Teams for successfully completion of major projects like Sensitive Data in Test, Production Data Protection - Sensitive Data Identification and Ownership.
- Adhere to HIPAA (Health Insurance Portability and Accountability Act) and data protection standards.
- Participating in global project planning and roadmap definition and Interacting with external partners, customers, and vendors.
- Performed Annual BCP (Business Continuity Plan) and Disaster recovery.
Application Developer
Confidential
Responsibilities:
- Responsibilities include attending Use Case Workshops, understanding doc preparation, estimation, developing, testing and implementation.
- Loading all flat files and DB2 tables data from various Applications to HDFS for further processing.
- Written and executed Apache PIG scripts on top of the HDFS data.
- Created Hive tables to store the processed results in a tabular format.
- Developed the Sqoop scripts in order to make the interaction between Pig and MySQL Database.
- Creating the script files for processing data and loading to HDFS.
- Writing CLI commands using HDFS.
- Developed the UNIX shell scripts for creating the reports from Hive data.
- Analyzing the requirement to setup a cluster.
- Setting up Cron job to delete Hadoop logs/local old job files/cluster temp files.
- Written Map Reduce code that will take input as log files and parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Effective coordination with Release teams during implementation.
- Prepare deployment plan for implementation projects.
- Fill out all documentation needed for production such as change management templates, instructions and deliverables.
- Coordinate PROD & TEST implementation and checkout.
- Share knowledge and experience with other team members during knowledge sharing sessions.
- Involved in System Design and Architecture of SAS Travel Agent Redemption Project, SAS Campaign Credits, NDP (New Data Project) to be imparted in System using IBM Mainframe Technologies.
- Coordinated and worked on critical IBM Mainframe Applications like Corporate Systems and Columbus Tracking Systems.
- Responsibilities include interacting with Business partners and customers in various phases of SDLC.
- Direct and lead the work of other Team members effectively.
Tech Lead and Senior Software Developer
Confidential, Jacksonville, FL
Responsibilities:
- Recommends, establishes business case, and assists project manager and others in building acceptance of new proposed program modifications, methods, and procedures. Conducts and provides moderately complex cost/benefit analysis for new proposed program modifications, methods, and procedures.
- Responsibilities include interacting with Business partners and customers in various phases of SDLC.
- Co-ordinate the work request in EDI, DARTS, TAXI, DARTS, Freight Claims and SMS area.
- Worked on Super Critical and business critical Transactions in EDI, TAXI.
- Development and Enhancements to the existing EDI and SMS system as specified by the customer using CICS, COBOL, Batch, DB2 & IMS
- Responsibilities include doing enhancements, maintenance, testing and assigning tasks to team members, guiding, conducting trainings and conducting project meetings.
- Interaction with Onsite counterpart and the users.
- Trained Team members and assigned the tasks and helping them to deliver the assigned tasks in time with no issues.
- Coordination of Production Releases, including major, minor, and hotfix releases.
- Prepared project related documents and WSR’s.
- Migration of OS/VS COBOL programs into enterprise COBOL.
- Converting, testing and delivering the converted programs and Reviewing the components to meet quality assurance.
