Job details
About the End Client and their Product & Services:
About the end client project: Give few lines about your client and the project.
Required Education: “Bachelor’s in Computers or relevant technology field ”.
Skills obtained from Education: “please specify some relevant courses in degree related to work/job duties”
Percentage of time spent on each duty: Please see near each duty.
Detailed Job duties:
Analyze the Hadoop Big Data distributed computing requirements from the system users and then design and develop the MapReduce and Spark components of the big data system to meet the user requirements. Understand the current business processes and leverage the Cloudera Manager and Navigator Web APIs by designing, developing and re-engineering Python and Unix/Linux shell-based software programs/tools for automating the manual tasks to improve the operational efficiency of the business. Analyze the performance of the Big Data system, find the bottlenecks in the system and then fine tune and customize the system for optimal performance of the business-critical applications. These tasks require understanding and practice of complex software engineering principles. 25%
Model and review the database local and external table designs in collaboration with other team members for optimal performance. Identify and implement optimal compression codecs and Hadoop file formats from storage and computation perspective. Compute the table statistics and refresh table metadata. 10%
Development and tuning of Hadoop based data warehouse. Benchmark applications to allocate required resources to mappers and reducers of the MapReduce and plan & design the cluster capacity to meet the business demands. Monitor resource manager, node manager services and application master for optimal performance and container level tuning. 15%
Implement security best practices through authentication and authorization of users and groups to cluster resources through LDAP protocols and Kerberos. Use Apache Sentry for role based fine grain authorization at table and column level for regulatory compliance readiness. 10%
Research and analyze the latest features and security vulnerability fixes in the release notes of the Oracle & Redhat Unix/Linux Operating systems and MySQL relational databases , recommend the administrative maintenance of the systems(Install/Upgrade) running those software components and define Key Performance Indicators (KPIs) for the system operation. Fine tune the systems to meet the required KPIs and design & implement backup and disaster recovery methods using proven best software practices like replication and mirroring for business continuity. 15%
Implement Cloudera distribution of Hadoop on Linux machines. Configure the HDFS storage and Name Node/Data Node services and make sure that all the components of the big data distributed system are working seamlessly together. Perform maintenance(Install/Upgrade) of the big data ecosystem tools and services and make sure that they are available Install, upgrade and maintain Hadoop ecosystem tools. 15%
As part of Life Cycle Management(LCM) of the Hadoop big data distributed computing platform, troubleshoot and solve issues related to Operating systems, databases and bigdata systems. Track production tickets in ticket monitoring tools, perform Root Cause Analysis(RCA) and resolve the tickets. 10%